Large Language Models Compared
The major models available to the public, side by side. Updated monthly so you can see how the field moves.
| Feature | πΊπΈGPT-5.6 SolOpenAI | πΊπΈClaude Opus 5Anthropic | πΊπΈGemini 3.1 ProGoogle DeepMind | πΊπΈGrok 4xAI | πΊπΈLlama 4Meta | π¨π³DeepSeek V4DeepSeek | π¨π³Qwen 3Alibaba Cloud | π«π·Mistral Large 3Mistral AI |
|---|---|---|---|---|---|---|---|---|
| Owner | OpenAI Group PBC, controlled by the non-profit OpenAI Foundation. Microsoft is the largest outside investor | Anthropic PBC. Independent, with Amazon and Google as major minority investors | Alphabet Inc, publicly listed. Founders retain voting control through dual-class shares | xAI Holdings. Elon Musk holds the majority economic stake. Saudi PIF, Sequoia and a16z are reported investors | Meta Platforms Inc, publicly listed. Mark Zuckerberg retains voting control | Owned by High-Flyer, a Chinese quantitative hedge fund based in Hangzhou | Alibaba Group Holding, publicly listed in Hong Kong and New York | Mistral AI SAS, French, founder-led. Microsoft holds a small minority stake |
| Context window | 1M tokens | 1M tokens | 2.5M tokens, the largest available | 1M tokens | 10M tokens on Scout variant | 128K to 1M tokens | 1M tokens | 256K tokens |
| Coding ability | ~84% SWE-bench | 87.6% to 95% SWE-bench | ~80.6% SWE-bench | ~72% SWE-bench | ~62% SWE-bench | 80.6% SWE-bench | ~70% SWE-bench | ~68% SWE-bench |
| Output speed | ~75 tok/sec | ~60 tok/sec | ~120 tok/sec, fastest of the flagships | ~90 tok/sec | Depends on your hosting | ~70 tok/sec | ~85 tok/sec | ~95 tok/sec |
| Price, USD per 1M tokens in / out | 0.20 / 1.60 | 5.00 / 25.00 | 1.25 / 10.00 | 3.00 / 15.00 | Free weights, you pay hosting | 0.14 / 0.28 | 0.30 / 1.20 | 2.00 / 6.00 |
| Openness | Closed weights | Closed weights | Closed weights | Partially open, older weights released | Open weights, community licence | Open weights on several variants | Open weights, Apache licence on most sizes | Open weights on smaller models |
| Reliability | Very high, mature API and enterprise tooling | Very high. Leads current intelligence and coding rankings | High. Deep integration across Google Workspace and Android | Moderate. Rapid release cycle, less enterprise track record | Good, but you own the operational burden | Good technically. Governance and data-residency questions for some buyers | Good. Strongest non-English performance in the group | Good. Positioned around European data sovereignty |
| Environmental disclosure | Not disclosed per model. Microsoft reports data-centre emissions at group level | Not disclosed per model | Google reports carbon-neutral operations claims at company level, not per model | Colossus data centre criticised locally over gas turbine emissions | Meta discloses training compute. Self-hosting shifts footprint to you | Notably efficient training. Lower compute cost per unit of capability | Not disclosed per model | Publishes more environmental detail than most peers |
| Adoption and niches | Broadest general adoption. Strong in customer service, marketing, education and general office work | Software engineering, legal, finance, regulated industries and long-document analysis | Anything document-heavy, plus consumer reach through Android and Search | Social media analysis, real-time information, consumer chat through X | Organisations needing on-premise control. Popular in defence, healthcare and government | Cost-sensitive deployments, high-volume automation, strong uptake across Asia | Chinese and multilingual markets, e-commerce, logistics. Growing Arabic capability | European public sector, defence and regulated industry seeking EU jurisdiction |
| Best suited to | Default choice when you want the widest ecosystem and integrations | Complex reasoning, code and work where being right matters more than being cheap | Very large documents, video and audio, and organisations already on Google Workspace | Real-time public discourse and anything needing current social data | Data that cannot leave your building, or heavy volume where per-token pricing hurts | Very high volume work where cost dominates and data sensitivity is low | Multilingual work and Asian market operations | When EU data residency or GDPR posture is the deciding factor |
What these differences actually mean
Price has stopped being the deciding factor for most work. The gap between the cheapest and most expensive model in this table is roughly thirty-fold on input tokens, yet for ordinary office tasks such as drafting, summarising and translating, the cheap models are entirely adequate. If cost was your reason for not adopting AI, that reason has expired. The real constraint now is whether your people know how to use these tools well, which is a training problem rather than a budget problem.
The expensive models earn their price in a narrow band of work. Complex reasoning, software engineering, legal and financial analysis, and any task where a confident wrong answer is costly. If your use case is drafting a customer email, paying premium rates buys you very little. If it is reviewing a contract or debugging production code, it buys you a great deal.
Ownership is becoming a procurement question, not a curiosity. Note how the ownership row splits: American models dominate general adoption, Chinese models compete aggressively on price and openness, and the single European entry sells almost entirely on jurisdiction. For organisations in Saudi Arabia this matters twice over, because data residency requirements and the Kingdom sovereign AI strategy both push towards asking where a model runs and who ultimately controls it. That question now belongs in vendor selection alongside price and capability.
Open weights change who carries the risk. An open model that you host yourself means your data never leaves your infrastructure, but it also means you own the security, the uptime and the cost of running it. Closed models reverse that trade. Neither is inherently safer, and the right answer depends on whether your constraint is confidentiality or capacity.
Environmental reporting in this industry is poor, and the table shows it. Almost no provider publishes per-model energy or emissions figures, offering company-level claims instead. That gap is itself worth noting: any organisation with sustainability commitments is currently unable to account accurately for the footprint of its AI usage. Expect pressure on this to grow.
The practical conclusion for most organisations is that model choice matters less than usage. The differences in this table are real but narrower than the marketing suggests, and they shift every few months. Consistently, the organisations getting value from these tools are not those who picked the highest-scoring model. They are the ones who taught their people what these tools do well, where they fail, and when to stop trusting the output.