Train with Julien

Researched and updated automatically through AI, monthly

Large Language Models Compared

The major models available to the public, side by side. Updated monthly so you can see how the field moves.

Last updated 9 August 2026
Scroll the table sideways to see all models
FeatureπŸ‡ΊπŸ‡ΈGPT-5.6 SolOpenAIπŸ‡ΊπŸ‡ΈClaude Opus 5AnthropicπŸ‡ΊπŸ‡ΈGemini 3.1 ProGoogle DeepMindπŸ‡ΊπŸ‡ΈGrok 4xAIπŸ‡ΊπŸ‡ΈLlama 4MetaπŸ‡¨πŸ‡³DeepSeek V4DeepSeekπŸ‡¨πŸ‡³Qwen 3Alibaba CloudπŸ‡«πŸ‡·Mistral Large 3Mistral AI
OwnerOpenAI Group PBC, controlled by the non-profit OpenAI Foundation. Microsoft is the largest outside investorAnthropic PBC. Independent, with Amazon and Google as major minority investorsAlphabet Inc, publicly listed. Founders retain voting control through dual-class sharesxAI Holdings. Elon Musk holds the majority economic stake. Saudi PIF, Sequoia and a16z are reported investorsMeta Platforms Inc, publicly listed. Mark Zuckerberg retains voting controlOwned by High-Flyer, a Chinese quantitative hedge fund based in HangzhouAlibaba Group Holding, publicly listed in Hong Kong and New YorkMistral AI SAS, French, founder-led. Microsoft holds a small minority stake
Context window1M tokens1M tokens2.5M tokens, the largest available1M tokens10M tokens on Scout variant128K to 1M tokens1M tokens256K tokens
Coding ability~84% SWE-bench87.6% to 95% SWE-bench~80.6% SWE-bench~72% SWE-bench~62% SWE-bench80.6% SWE-bench~70% SWE-bench~68% SWE-bench
Output speed~75 tok/sec~60 tok/sec~120 tok/sec, fastest of the flagships~90 tok/secDepends on your hosting~70 tok/sec~85 tok/sec~95 tok/sec
Price, USD per 1M tokens in / out0.20 / 1.605.00 / 25.001.25 / 10.003.00 / 15.00Free weights, you pay hosting0.14 / 0.280.30 / 1.202.00 / 6.00
OpennessClosed weightsClosed weightsClosed weightsPartially open, older weights releasedOpen weights, community licenceOpen weights on several variantsOpen weights, Apache licence on most sizesOpen weights on smaller models
ReliabilityVery high, mature API and enterprise toolingVery high. Leads current intelligence and coding rankingsHigh. Deep integration across Google Workspace and AndroidModerate. Rapid release cycle, less enterprise track recordGood, but you own the operational burdenGood technically. Governance and data-residency questions for some buyersGood. Strongest non-English performance in the groupGood. Positioned around European data sovereignty
Environmental disclosureNot disclosed per model. Microsoft reports data-centre emissions at group levelNot disclosed per modelGoogle reports carbon-neutral operations claims at company level, not per modelColossus data centre criticised locally over gas turbine emissionsMeta discloses training compute. Self-hosting shifts footprint to youNotably efficient training. Lower compute cost per unit of capabilityNot disclosed per modelPublishes more environmental detail than most peers
Adoption and nichesBroadest general adoption. Strong in customer service, marketing, education and general office workSoftware engineering, legal, finance, regulated industries and long-document analysisAnything document-heavy, plus consumer reach through Android and SearchSocial media analysis, real-time information, consumer chat through XOrganisations needing on-premise control. Popular in defence, healthcare and governmentCost-sensitive deployments, high-volume automation, strong uptake across AsiaChinese and multilingual markets, e-commerce, logistics. Growing Arabic capabilityEuropean public sector, defence and regulated industry seeking EU jurisdiction
Best suited toDefault choice when you want the widest ecosystem and integrationsComplex reasoning, code and work where being right matters more than being cheapVery large documents, video and audio, and organisations already on Google WorkspaceReal-time public discourse and anything needing current social dataData that cannot leave your building, or heavy volume where per-token pricing hurtsVery high volume work where cost dominates and data sensitivity is lowMultilingual work and Asian market operationsWhen EU data residency or GDPR posture is the deciding factor

What these differences actually mean

Price has stopped being the deciding factor for most work. The gap between the cheapest and most expensive model in this table is roughly thirty-fold on input tokens, yet for ordinary office tasks such as drafting, summarising and translating, the cheap models are entirely adequate. If cost was your reason for not adopting AI, that reason has expired. The real constraint now is whether your people know how to use these tools well, which is a training problem rather than a budget problem.

The expensive models earn their price in a narrow band of work. Complex reasoning, software engineering, legal and financial analysis, and any task where a confident wrong answer is costly. If your use case is drafting a customer email, paying premium rates buys you very little. If it is reviewing a contract or debugging production code, it buys you a great deal.

Ownership is becoming a procurement question, not a curiosity. Note how the ownership row splits: American models dominate general adoption, Chinese models compete aggressively on price and openness, and the single European entry sells almost entirely on jurisdiction. For organisations in Saudi Arabia this matters twice over, because data residency requirements and the Kingdom sovereign AI strategy both push towards asking where a model runs and who ultimately controls it. That question now belongs in vendor selection alongside price and capability.

Open weights change who carries the risk. An open model that you host yourself means your data never leaves your infrastructure, but it also means you own the security, the uptime and the cost of running it. Closed models reverse that trade. Neither is inherently safer, and the right answer depends on whether your constraint is confidentiality or capacity.

Environmental reporting in this industry is poor, and the table shows it. Almost no provider publishes per-model energy or emissions figures, offering company-level claims instead. That gap is itself worth noting: any organisation with sustainability commitments is currently unable to account accurately for the footprint of its AI usage. Expect pressure on this to grow.

The practical conclusion for most organisations is that model choice matters less than usage. The differences in this table are real but narrower than the marketing suggests, and they shift every few months. Consistently, the organisations getting value from these tools are not those who picked the highest-scoring model. They are the ones who taught their people what these tools do well, where they fail, and when to stop trusting the output.

Figures are approximate and gathered from public benchmarks and provider pricing pages. Benchmark scores vary by testing method. Prices exclude volume discounts and change frequently.