LLM performance rankings
Benchmark scores and performance-per-dollar across 110 benchmarked models. Updated with the latest public evaluations — every ranking pairs the score with what the model actually costs to run.
Models benchmarked
110
of 231 tracked models — scores from public leaderboards. One model counts once, however many hosts sell it.
Benchmarks tracked
16
GPQA Diamond · SWE-Bench Verified · Terminal-Bench 2.1 · LiveCodeBench · MCP Atlas · BrowseComp · OSWorld-Verified · Humanity's Last Exam · ARC-AGI 2 · AIME 2025 · MMMLU · BFCL · HumanEval · MATH 500 · AutoBench · AIME 2026
Cost vs quality — the value frontier
every dot is a model page — hover for the numbers50 models plotted by blended cost (log scale) against average benchmark score — hover any dot for details, click to open the model. The dashed line is the value frontier: no model is both cheaper and better than a point on it. Cheapest model scoring 80+: GPT-5.6 Luna at $0.450/Mtok. Only models with ≥3 of the 9 tracked benchmarks are averaged (21 thinly-tested models excluded — a one-benchmark "average" isn't comparable).
Best value — performance per dollar
norm value drives the ranking — top 25 shownRanked by comparability-normalized value — each model's benchmark strength is scored as a percentile rank against the whole tracked population, then divided by blended cost ($/Mtok), so a model tested only on easy benchmarks can't buy its way to the top. The Perf/$ column shows the raw ratio for reference; the Norm value column is what drives the ranking. Ranked models have ≥3 of the 9 tracked benchmarks — the "n" column shows how many; averages over one or two benchmarks aren't comparable and are excluded.
| Model | Norm value | Avg score (n) | Blended $/M | Perf/$ | Value |
|---|---|---|---|---|---|
| 1DeepSeek V4 FlashDeepInfra +2 more hosts | 515.5 | 77.5% (9) | from $0.110 | 688.9 | |
| 2GLM-5.3-FlashZ.AI | 343.8 | 69.8% (6) | $0.240 | 293.9 | |
| 3Solar Pro 4Upstage | 265.6 | 57% (3) | $0.160 | 361.9 | |
| 4GPT-OSS 20BFireworks +1 more host | 223.8 | 62.5% (6) | from $0.130 | 490.2 | |
| 5GPT-5.6 LunaOpenAI | 150.2 | 90.7% (8) | $0.450 | 201.6 | |
| 6GPT-OSS 120BTogether +2 more hosts | 136.9 | 65.5% (6) | from $0.260 | 249.5 | |
| 7Llama 4 ScoutDeepInfra +2 more hosts | 128 | 46.8% (12) | from $0.150 | 312 | |
| 8DeepSeek V4.1 FlashDeepSeek | 119 | 56.8% (5) | $0.520 | 108.2 | |
| 9MiniMax-M3MiniMax +2 more hosts | 83.8 | 63.1% (17) | from $0.520 | 120.2 | |
| 10Llama 4 MaverickDeepInfra | 79.4 | 65.1% (5) | $0.350 | 186 | |
| 11Llama 3.1 8BMeta | 70 | 34% (9) | $0.060 | 591.3 | |
| 12Llama 3.1 8B InstantGroq | 70 | 34% (9) | $0.060 | 591.3 | |
| 13Gemini 3.7 FlashGoogle | 53.1 | 85.8% (5) | $1.50 | 57.2 | |
| 14Muse Spark 1.3Meta | 44.6 | 67.5% (6) | $2.00 | 33.8 | |
| 15Mistral Large 3Mistral | 44.5 | 53.6% (12) | $0.750 | 71.5 | |
| 16Grok 4.3xAI | 41.8 | 72.4% (12) | $1.56 | 46.3 | |
| 17Muse Glimmer 30BFireworks +1 more host | 41.2 | 37% (6) | from $0.640 | 58 | |
| 18Qwen3.8-27BAlibaba | 40.8 | 34% (5) | $1.12 | 30.2 | |
| 19GPT-4.1 nanoOpenAI | 39.4 | 58.6% (3) | $0.180 | 334.9 | |
| 20Muse Spark 1.2Meta | 38.8 | 62% (6) | $2.00 | 31 | |
| 21GLM-5.3Z.AI +2 more hosts | 37.7 | 42% (5) | from $2.15 | 19.5 | |
| 22Claude Haiku 3.5Anthropic | 36.7 | 78.8% (3) | $1.60 | 49.2 | |
| 23DeepSeek V4 ProDeepInfra +3 more hosts | 36.2 | 73.5% (14) | from $1.62 | 45.2 | |
| 24Gemini 2.5 FlashGoogle | 35.3 | 50.2% (10) | $0.850 | 59.1 | |
| 25Kimi K2.6Fireworks +1 more host | 35.1 | 76.2% (7) | from $1.71 | 44.5 |
Composite intelligence scores
published values only — two scales, never added togetherTwo independent, cross-model composite indices. Unlike the per-benchmark tables below, these summarise overall capability in a single number — and for the newest family flagships (which the per-benchmark leaderboards don't yet cover) they are often the only real, independent scores that exist. We show published values only; a — means no composite has been released for that model. Sources: Artificial Analysis Intelligence Index and LMSYS Chatbot Arena ELO (UC Berkeley).
Artificial Analysis Intelligence Index
A 0–100 blended index across reasoning, math, coding and knowledge evals. Higher is better.
| Model | AA index | Cost $/M |
|---|---|---|
| 1Claude Fable 5.1Anthropic | 53.4 | $20.00 |
| 2GPT-6 AstraOpenAI | 52.8 | $20.00 |
| 3DeepSeek V4 Flash Vision ExpDeepSeek | 51 | $0.660 |
| 4Claude Opus 5Anthropic | 50.7 | $10.00 |
| 5Claude Fable 5Anthropic | 49.7 | $20.00 |
| 6Muse Spark 1.3Meta | 48.2 | $2.00 |
| 7GPT-5.6 SolOpenAI | 47.1 | $8.00 |
| 8Kimi K2.6Fireworks | 45.1 | $1.71 |
| 9GLM-5.3Z.AI +2 more hosts | 44.9 | from $2.15 |
| 10Grok 4.6xAI | 44.4 | $3.00 |
| 11GPT-5.3 CodexOpenAI | 44.3 | $4.81 |
| 12Kimi K3Moonshot | 43.8 | $6.00 |
| 13GPT-5.6 TerraOpenAI | 42.3 | $4.50 |
| 14GPT-5.2OpenAI | 42.2 | $4.81 |
| 15Claude Opus 4.8Anthropic | 42 | $10.00 |
| 16Solar Pro 4Upstage | 42 | $0.160 |
| 17GLM-5.3-FlashZ.AI | 41.9 | $0.240 |
| 18Gemini 3.8 FlashGoogle | 41.2 | $1.50 |
| 19Qwen3.6-PlusFireworks | 40.5 | $1.12 |
| 20Qwen3.8-MaxAlibaba | 40.3 | $3.00 |
LMSYS Arena ELO
Crowd head-to-head preference rating from blind human votes. Typically 1300–1550; higher is better.
| Model | Arena ELO | Cost $/M |
|---|---|---|
| 1GPT-5.5 ProOpenAI | 1510 | $67.50 |
| 2Claude Fable 5Anthropic | 1506 | $20.00 |
| 3Muse Spark 1.2Meta | 1500 | $2.00 |
| 4Claude Fable 5.1Anthropic | 1499 | $20.00 |
| 5Claude Opus 4.6Anthropic | 1497 | $10.00 |
| 6Claude Opus 4.7Anthropic | 1494 | $10.00 |
| 7Muse Spark 1.3Meta | 1493 | $2.00 |
| 8Gemini 3.8 FlashGoogle | 1493 | $1.50 |
| 9Claude Opus 5Anthropic | 1493 | $10.00 |
| 10Muse Spark 1.1Meta | 1493 | $2.00 |
| 11Gemini 3.7 FlashGoogle | 1490 | $1.50 |
| 12Gemini 3.1 ProGoogle | 1487 | $4.50 |
| 13Kimi K3Moonshot | 1485 | $6.00 |
| 14GPT-5.6 SolOpenAI | 1484 | $8.00 |
| 15GLM-5.3Z.AI | 1483 | $2.15 |
| 16Claude Opus 4.8Anthropic | 1481 | $10.00 |
| 17Qwen3.8-MaxAlibaba | 1481 | $3.00 |
| 18Gemini 3.6 FlashGoogle | 1480 | $1.50 |
| 19GPT-6 AstraOpenAI | 1480 | $20.00 |
| 20GPT-5.4 ProOpenAI | 1478 | $67.50 |
These two indices use different scales and can't be added together — a model can rank highly on one and be absent from the other. Composite scores complement, and don't replace, the per-benchmark accuracy tables below.
The per-benchmark boards below rank independently measured scores. A maker's own number is flagged "vendor", listed after the independent rows, and never competes for a rank.
Best reasoning — GPQA Diamond
PhD-level science reasoning. The hardest general reasoning benchmark.
| Model | GPQA score | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Sonnet 5Anthropic | 96.2% | $4.00 | 20.4 |
| 2GPT-6 AstraOpenAI | 96% | $20.00 | 4.2 |
| 3GPT-5.6 SolOpenAI | 94.6% | $8.00 | 10.5 |
| 4Gemini 3.1 ProGoogle | 94.3% | $4.50 | 16.6 |
| 5Claude Opus 4.7Anthropic | 94.2% | $10.00 | 8.3 |
| 6Claude Mythos 5Anthropic | 94.1% | $20.00 | 4.4 |
| 7Claude Fable 5Anthropic | 94.1% | $20.00 | 4.5 |
| 8Claude Opus 4.8Anthropic | 93.6% | $10.00 | 8.1 |
| 9GPT-5.5OpenAI | 93.6% | $11.25 | 6.7 |
| 10Kimi K3Moonshot | 93.5% | $6.00 | 13.9 |
| 11Claude Fable 5.1Anthropic | 93.4% | $20.00 | 4.2 |
| 12MiniMax-M3MiniMax +2 more hosts | 93% | from $0.520 | 120.2 |
| 13GPT-5.6 TerraOpenAI | 92.9% | $4.50 | 19.5 |
| 14GPT-5.2OpenAI | 92.4% | $4.81 | 15.8 |
| 15GPT-5.6 LunaOpenAI | 92.3% | $0.450 | 201.6 |
Best coding — SWE-Bench
Real-world software engineering tasks — resolving GitHub issues across popular repos.
| Model | SWE-Bench | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Opus 5Anthropic | 97% | $10.00 | 8.4 |
| 2GPT-5.6 SolOpenAI | 96.2% | $8.00 | 10.5 |
| 3Claude Mythos 5Anthropic | 95.5% | $20.00 | 4.4 |
| 4Claude Fable 5Anthropic | 95% | $20.00 | 4.5 |
| 5Kimi K3Moonshot | 93.4% | $6.00 | 13.9 |
| 6GPT-5.6 LunaOpenAI | 93% | $0.450 | 201.6 |
| 7Claude Opus 4.8Anthropic | 88.6% | $10.00 | 8.1 |
| 8Claude Opus 4.7Anthropic | 87.6% | $10.00 | 8.3 |
| 9Claude Sonnet 5Anthropic | 85.2% | $4.00 | 20.4 |
| 10Claude Sonnet 4.5Anthropic | 82% | $6.00 | 12.5 |
| 11GPT-5.4OpenAI | 75% | $5.62 | 14 |
| 12o4-miniOpenAI | 68% | $1.93 | 35.9 |
| 13Grok 4.3xAI | 65% | $1.56 | 46.3 |
| 14Grok 4xAI | 55% | $6.00 | 10.5 |
| 15Gemini 2.5 ProGoogle | 50% | $3.44 | 17.8 |
Terminal-Bench 2.1
Multi-step tasks in a terminal environment.
| Model | Terminal-Bench 2.1 | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Fable 5.1Anthropic | 91.4% | $20.00 | 4.2 |
| 2GPT-5.6 SolOpenAI | 88.8% | $8.00 | 10.5 |
| 3Kimi K3Moonshot | 88.3% | $6.00 | 13.9 |
| 4Claude Mythos 5Anthropic | 88% | $20.00 | 4.4 |
| 5Muse Spark 1.3Meta | 86% | $2.00 | 33.8 |
| 6Gemini 3.7 FlashGoogle | 85.8% | $1.50 | 57.2 |
| 7Claude Fable 5Anthropic | 84.3% | $20.00 | 4.5 |
| 8GPT-5.6 LunaOpenAI | 84.3% | $0.450 | 201.6 |
| 9GLM-5.3-FlashZ.AI | 84.3% | $0.240 | 293.9 |
| 10GPT-5.5OpenAI | 82.7% | $11.25 | 6.7 |
| 11GPT-5.6 TerraOpenAI | 82.5% | $4.50 | 19.5 |
| 12GLM-5.2Z.AI +2 more hosts | 81% | from $2.15 | 35.3 |
| 13Claude Sonnet 5Anthropic | 80.4% | $4.00 | 20.4 |
| 14Muse Spark 1.2Meta | 80% | $2.00 | 31 |
| 15Muse Spark 1.1Meta | 78% | $2.00 | 30.8 |
LiveCodeBench
Contamination-resistant code generation.
| Model | LiveCodeBench | Cost $/M | Perf/$ |
|---|---|---|---|
| 1DeepSeek V4 ProDeepSeek +3 more hosts | 93.5% | from $1.62 | 37.1 |
| 2DeepSeek V4 FlashDeepSeek +2 more hosts | 91.6% | from $0.110 | 117.4 |
| 3Claude Fable 5.1Anthropic | 90.5% | $20.00 | 4.2 |
| 4Kimi K2.5Moonshot +1 more host | 85% | from $1.20 | 52.4 |
| 5Grok 4xAI | 79% | $6.00 | 10.5 |
| 6Claude Opus 4.6Anthropic | 76% | $10.00 | 7.7 |
| 7o3-miniOpenAI | 74.1% | $1.93 | 34.2 |
| 8Claude Sonnet 4.6Anthropic | 72.4% | $6.00 | 11.7 |
| 9Gemini 2.5 ProGoogle | 69% | $3.44 | 17.8 |
| 10GPT-OSS 20BFireworks +1 more host | 69% | from $0.130 | 490.2 |
| 11GPT-OSS 120BFireworks +2 more hosts | 69% | from $0.260 | 249.5 |
| 12Gemini 2.5 FlashGoogle | 63.5% | $0.850 | 59.1 |
| 13GPT-4.1OpenAI | 52% | $3.50 | 16.3 |
| 14Llama 4 MaverickDeepInfra | 41% | $0.350 | 186 |
| 15Llama 4 ScoutMeta +2 more hosts | 32.8% | from $0.150 | 279.4 |
MCP Atlas
Tool use over Model Context Protocol servers.
| Model | MCP Atlas | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Kimi K3Moonshot | 84.2% | $6.00 | 13.9 |
| 2Gemini 3.5 FlashGoogle | 83.6% | $3.38 | 20.8 |
| 3Claude Fable 5Anthropic | 83.3% | $20.00 | 4.5 |
| 4Claude Opus 4.8Anthropic | 82.2% | $10.00 | 8.1 |
| 5Claude Opus 4.7Anthropic | 79.1% | $10.00 | 8.3 |
| 6GLM-5.2Z.AI +2 more hosts | 77% | from $2.15 | 35.3 |
| 7Claude Opus 4.6Anthropic | 76.8% | $10.00 | 7.7 |
| 8GPT-5.5OpenAI | 75.3% | $11.25 | 6.7 |
| 9MiniMax-M3MiniMax +2 more hosts | 74.2% | from $0.520 | 120.2 |
| 10DeepSeek V4 ProDeepSeek +3 more hosts | 73.6% | from $1.62 | 37.1 |
| 11Claude Opus 4.5Anthropic | 69.8% | $10.00 | 7.1 |
| 12Claude Sonnet 4.6Anthropic | 69.5% | $6.00 | 11.7 |
| 13Gemini 3.1 ProGoogle | 69.2% | $4.50 | 16.6 |
| 14DeepSeek V4 FlashDeepSeek +2 more hosts | 69% | from $0.110 | 117.4 |
| 15GPT-5.2OpenAI | 67.6% | $4.81 | 15.8 |
BrowseComp
Agentic web search and information extraction.
| Model | BrowseComp | Cost $/M | Perf/$ |
|---|---|---|---|
| 1GPT-5.6 SolOpenAI | 92.2% | $8.00 | 10.5 |
| 2GPT-6 AstraOpenAI | 91.5% | $20.00 | 4.2 |
| 3Kimi K3Moonshot | 91.2% | $6.00 | 13.9 |
| 4Claude Opus 5Anthropic | 90.8% | $10.00 | 8.4 |
| 5Claude Fable 5Anthropic | 88% | $20.00 | 4.5 |
| 6Gemini 3.1 ProGoogle | 85.9% | $4.50 | 16.6 |
| 7DeepSeek V4 FlashDeepSeek +2 more hosts | 85.9% | from $0.110 | 117.4 |
| 8Claude Sonnet 5Anthropic | 84.7% | $4.00 | 20.4 |
| 9GPT-5.5OpenAI | 84.4% | $11.25 | 6.7 |
| 10MiniMax-M3MiniMax +2 more hosts | 83.5% | from $0.520 | 120.2 |
| 11DeepSeek V4 ProDeepSeek +3 more hosts | 83.4% | from $1.62 | 37.1 |
| 12Kimi K2.6Moonshot +1 more host | 83.2% | from $1.71 | 44.5 |
| 13Claude Sonnet 4.6Anthropic | 76.2% | $6.00 | 11.7 |
| 14Inkling vendorThinking Machines | 77.1% | $1.76 | — |
| 15Solar Pro 4 vendorUpstage | 49.2% | $0.160 | 361.9 |
OSWorld-Verified
Computer use — real GUI tasks on a desktop OS.
| Model | OSWorld-Verified | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Fable 5Anthropic | 85% | $20.00 | 4.5 |
| 2Claude Opus 4.8Anthropic | 83.4% | $10.00 | 8.1 |
| 3Claude Sonnet 5Anthropic | 81.2% | $4.00 | 20.4 |
| 4GPT-5.5OpenAI | 78.7% | $11.25 | 6.7 |
| 5Claude Sonnet 4.6Anthropic | 78.5% | $6.00 | 11.7 |
| 6Gemini 3.5 FlashGoogle | 78.4% | $3.38 | 20.8 |
| 7Claude Opus 4.7Anthropic | 78% | $10.00 | 8.3 |
| 8Gemini 3.1 ProGoogle | 76.2% | $4.50 | 16.6 |
| 9Kimi K2.6Moonshot +1 more host | 73.1% | from $1.71 | 44.5 |
| 10MiniMax-M3MiniMax +2 more hosts | 70.1% | from $0.520 | 120.2 |
| 11GPT-5.3 CodexOpenAI | 64.7% | $4.81 | 14.8 |
| 12GPT-5.6 SolOpenAI | 62.6% | $8.00 | 10.5 |
| 13Qwen3.8-27B vendorAlibaba | 84.3% | $1.12 | 30.2 |
| 14Gemini 3.6 Flash vendorGoogle | 83% | $1.50 | — |
| 15Muse Spark 1.1 vendorMeta | 80.8% | $2.00 | 30.8 |
Best math — AIME 2025
American Invitational Mathematics Examination problems — competition-level math.
| Model | AIME score | Cost $/M | Perf/$ |
|---|---|---|---|
| 1GPT-5.2OpenAI | 100% | $4.81 | 15.8 |
| 2Claude Opus 4.6Anthropic | 99.8% | $10.00 | 7.7 |
| 3GPT-OSS 20BFireworks +1 more host | 98.7% | from $0.130 | 490.2 |
| 4GPT-OSS 120BFireworks +2 more hosts | 97.9% | from $0.260 | 249.5 |
| 5Claude Haiku 4.5Anthropic | 96.3% | $2.00 | 32.8 |
| 6Kimi K2.5Moonshot +1 more host | 96.1% | from $1.20 | 52.4 |
| 7GPT-5.4OpenAI | 95.5% | $5.62 | 14 |
| 8o4-miniOpenAI | 92.7% | $1.93 | 35.9 |
| 9Grok 4xAI | 91.7% | $6.00 | 10.5 |
| 10Gemini 2.5 ProGoogle | 88% | $3.44 | 17.8 |
| 11Claude Sonnet 4.5Anthropic | 87% | $6.00 | 12.5 |
| 12o3-miniOpenAI | 86.5% | $1.93 | 34.2 |
| 13Grok 4.3xAI | 85% | $1.56 | 46.3 |
| 14Claude Sonnet 4.6Anthropic | 83% | $6.00 | 11.7 |
| 15Claude Opus 4.1Anthropic | 78% | $30.00 | 2.1 |
Fastest endpoints — inference speed
Tokens per second on hosted inference — the one board here ranked by endpoint rather than by model, because speed is a property of the host's serving stack, not of the weights. The same model can appear more than once at different hosts, and the difference between those rows is the point. Speed matters for real-time applications and high-throughput pipelines.
| Model | Tokens/sec | Cost $/M |
|---|---|---|
| 1Llama 4 ScoutMeta | 2600 t/s | $0.170 |
| 2Llama 4 ScoutGroq | 2600 t/s | $0.170 |
| 3Llama 4 ScoutDeepInfra | 2600 t/s | $0.150 |
| 4Llama 3.3 70BMeta | 2500 t/s | $0.640 |
| 5Llama 3.3 70B VersatileGroq | 2500 t/s | $0.640 |
| 6Llama 3.3 70BTogether | 2500 t/s | $1.04 |
| 7GPT-OSS 120BFireworks | 1828.8 t/s | $0.260 |
| 8Llama 3.1 8BMeta | 1800 t/s | $0.060 |
| 9Llama 3.1 8B InstantGroq | 1800 t/s | $0.060 |
| 10GPT-OSS 20BFireworks | 941.5 t/s | $0.130 |
| 11Gemini 3.1 Flash-LiteGoogle | 600 t/s | $0.560 |
| 12GPT-OSS 20BGroq | 564 t/s | $0.130 |
| 13Qwen3.6-27BGroq | 472.4 t/s | $1.20 |
| 14Gemini 2.5 Flash-LiteGoogle | 450 t/s | $0.180 |
| 15Granite 4 H SmallIBM | 414.5 t/s | $0.110 |
Hardest benchmark — Humanity's Last Exam
Expert-level questions across 100+ fields. No model scores above 65% — the frontier of AI capability.
| Model | HLE score | Cost $/M |
|---|---|---|
| 1Claude Opus 5Anthropic | 64.7% | $10.00 |
| 2Claude Mythos 5Anthropic | 64.5% | $20.00 |
| 3Claude Fable 5.1Anthropic | 59.1% | $20.00 |
| 4Claude Opus 4.8Anthropic | 57.9% | $10.00 |
| 5Claude Sonnet 5Anthropic | 57.4% | $4.00 |
| 6GPT-6 AstraOpenAI | 57.2% | $20.00 |
| 7Kimi K3Moonshot | 56% | $6.00 |
| 8GLM-5.3-FlashZ.AI | 55.3% | $0.240 |
| 9GLM-5.2Z.AI +2 more hosts | 54.7% | from $2.15 |
| 10Kimi K2.6Moonshot +1 more host | 54% | from $1.71 |
Methodology
how the scores and rankings are computedBenchmark sources
Vellum LLM Leaderboard, Artificial Analysis, official model technical reports, and HuggingFace Open LLM Leaderboard. Scores are the latest publicly reported results as of September 2026, refreshed as new evaluations are published.
Performance per dollar
Shown two ways. The raw Perf/$ column is avg_benchmark_score / blended_cost_per_mtok (blended cost = (3×input + 1×output) ÷ 4 per million tokens). The Best Value ranking is driven instead by the comparability-normalized Norm Value: each model's benchmark strength is first converted to a percentile rank against the whole tracked population (avg_method = percentile_rank_vs_population), then divided by the same blended cost — so a model tested only on easy benchmarks can't post an inflated ratio.
A Norm Value is a percentile-derived value figure, not a % accuracy — only the raw Avg Score and per-benchmark columns are accuracy percentages.
Comparability rule
A model's "average" only covers the benchmarks it has actually been tested on, and the tracked benchmarks differ wildly in difficulty (nobody scores above ~65% on Humanity's Last Exam; AIME scores run into the high 90s), so averaging over different subsets isn't apples-to-apples. Every avg-based ranking and the scatter above therefore require at least 3 filled benchmarks and display the benchmark count (n). Single-benchmark tables (GPQA, SWE-Bench, Terminal-Bench, AIME, HLE and the rest) are unaffected — one shared benchmark is directly comparable.
One row per model
Open-weight models are often sold by several hosts at once — GLM-5.2 by Z.AI, Fireworks and Together. A benchmark score belongs to the weights, so all three score identically, and listing them separately would spend three ranking slots on one answer. Every board here shows a model once, at its cheapest host, with the remaining hosts noted on the row; the price reads “from” whenever more than one host sells it. The counts above follow the same rule — a model sold by three hosts counts once.
The one exception is Fastest endpoints. Tokens per second is a property of the host's serving stack rather than of the weights — Z.AI runs GLM-5.2 at a different speed from Fireworks — so that board ranks endpoints, and a model can legitimately appear on it more than once.
What each benchmark measures
- GPQA Diamond
- PhD-level questions in biology, chemistry, and physics
- AIME 2025
- American Invitational Mathematics Examination (competition math)
- SWE-Bench
- Resolving real GitHub issues in popular Python repositories
- Humanity's Last Exam
- Expert-level questions across 100+ academic fields
- ARC-AGI 2
- Visual reasoning and abstract pattern completion
- MMMLU
- Multilingual version of MMLU (57 subjects, multiple languages)
- HumanEval
- Python code generation from natural language descriptions
- MATH 500
- Mathematical problem solving across 5 difficulty levels
- BFCL
- Berkeley Function Calling Leaderboard (tool use & function calling)
Access this data programmatically via the benchmarks API endpoint or the API documentation.