ModelPriceWatch.com
Last scan 2026-08-09 Models tracked 198 Providers 30 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

LLM performance rankings

Benchmark scores and performance-per-dollar across 89 benchmarked models. Updated with the latest public evaluations — every ranking pairs the score with what the model actually costs to run.

Models benchmarked

89

of 166 tracked models — scores from public leaderboards. One model counts once, however many hosts sell it.

Benchmarks tracked

9

GPQA · AIME · SWE-Bench · HLE · ARC-AGI 2 · MMMLU · HumanEval · MATH 500 · BFCL

Providers covered

23

each score keeps its source — how we count

Cost vs quality — the value frontier

every dot is a model page — hover for the numbers

37 models plotted by blended cost (log scale) against average benchmark score — hover any dot for details, click to open the model. The dashed line is the value frontier: no model is both cheaper and better than a point on it. Cheapest model scoring 80+: GPT-5.6 Luna at $0.450/Mtok. Only models with ≥3 of the 9 tracked benchmarks are averaged (19 thinly-tested models excluded — a one-benchmark "average" isn't comparable).

30 40 50 60 70 80 90 $0.05 $0.1 $0.3 $1 $3 $10 Blended cost $/Mtok (log scale) → Avg benchmark score → DeepSeek V4 Flash GPT-5.6 Luna Llama 3.1 8B
Closed weights Open weights Best value Value frontier MODELPRICEWATCH.COM · 2026-08-09

Best value — performance per dollar

norm value drives the ranking — top 25 shown

Ranked by comparability-normalized value — each model's benchmark strength is scored as a percentile rank against the whole tracked population, then divided by blended cost ($/Mtok), so a model tested only on easy benchmarks can't buy its way to the top. The Perf/$ column shows the raw ratio for reference; the Norm value column is what drives the ranking. Ranked models have ≥3 of the 9 tracked benchmarks — the "n" column shows how many; averages over one or two benchmarks aren't comparable and are excluded.

Best-value LLM ranking — benchmark strength per blended dollar, top 25 of 89 benchmarked models
Model Norm value Avg score (n) Blended $/M Perf/$ Value
1DeepSeek V4 FlashDeepInfra +2 more hosts 568.2 77.5% (10) from $0.110 688.9
2GPT-OSS 20BFireworks +1 more host 223.8 62.5% (7) from $0.130 490.2
3GPT-5.6 LunaOpenAI 174.2 90.7% (8) $0.450 201.6
4GPT-OSS 120BTogether +2 more hosts 138.8 65.5% (7) from $0.260 249.5
5Llama 4 ScoutDeepInfra +2 more hosts 126 46.8% (13) from $0.150 312
6DeepSeek V4 ProDeepSeek +2 more hosts 108.3 73.5% (16) from $0.540 135.2
7MiniMax-M3MiniMax +2 more hosts 93.8 63.1% (18) from $0.520 120.2
8Llama 3.1 8BMeta 73.3 34% (9) $0.060 591.3
9Llama 3.1 8B InstantGroq 73.3 34% (9) $0.060 591.3
10Llama 4 MaverickDeepInfra 71.1 65.1% (6) $0.350 186
11Mistral Large 3Mistral 45.7 53.6% (13) $0.750 71.5
12Grok 4.3xAI 41.1 72.4% (13) $1.56 46.3
13GPT-4.1 nanoOpenAI 37.8 58.6% (5) $0.180 334.9
14Claude Haiku 3.5Anthropic 37.4 78.8% (3) $1.60 49.2
15Gemini 2.5 FlashGoogle 37.1 50.2% (10) $0.850 59.1
16Kimi K2.5Moonshot +1 more host 35.7 62.9% (12) from $1.20 52.4
17GLM-5.2Z.AI +2 more hosts 35.1 76% (9) from $2.15 35.3
18Kimi K2.6Moonshot +1 more host 34.7 76.2% (9) from $1.71 44.5
19Llama 3.3 70BMeta +2 more hosts 34.5 46.6% (9) from $0.640 72.8
20o4-miniOpenAI 29.2 69.2% (10) $1.93 35.9
21GPT-4.1 miniOpenAI 25 55.7% (7) $0.700 79.6
22Gemini 3.5 FlashGoogle 22.9 70.1% (9) $3.38 20.8
23Claude Sonnet 5Anthropic 20.4 81.5% (11) $4.00 20.4
24GPT-5.6 TerraOpenAI 17.7 87.7% (6) $4.50 19.5
25Claude Haiku 4.5Anthropic 17.4 65.5% (10) $2.00 32.8
Norm value = percentile-rank benchmark strength ÷ blended $/Mtok. How we rank MODELPRICEWATCH.COM · 2026-08-09

Composite intelligence scores

published values only — two scales, never added together

Two independent, cross-model composite indices. Unlike the per-benchmark tables below, these summarise overall capability in a single number — and for the newest family flagships (which the per-benchmark leaderboards don't yet cover) they are often the only real, independent scores that exist. We show published values only; a means no composite has been released for that model. Sources: Artificial Analysis Intelligence Index and LMSYS Chatbot Arena ELO (UC Berkeley).

Artificial Analysis Intelligence Index

A 0–100 blended index across reasoning, math, coding and knowledge evals. Higher is better.

Top 20 models by Artificial Analysis Intelligence Index, with blended cost
ModelAA indexCost $/M
1Claude Opus 5Anthropic 63.1 $10.00
2Claude Fable 5Anthropic 62.1 $20.00
3GPT-5.6 SolOpenAI 60.9 $11.25
4Kimi K3Moonshot 59.7 $6.00
5Qwen3.8-MaxAlibaba 58.1 $3.00
6Claude Opus 4.8Anthropic 57.3 $10.00
7Muse Spark 1.2Meta 56.8 $2.00
8GPT-5.6 TerraOpenAI 56.6 $4.50
9GPT-5.5OpenAI 56.3 $11.25
10Grok 4.5xAI 55.8 $3.00
11Claude Sonnet 5Anthropic 55.3 $4.00
12Claude Opus 4.7Anthropic 55 $10.00
13Muse Spark 1.1Meta 53.2 $2.00
14GPT-5.4OpenAI 53.1 $5.62
15GLM-5.2Z.AI +2 more hosts 52.6 from $2.15
16GPT-5.6 LunaOpenAI 52.3 $0.450
17Gemini 3.5 FlashGoogle 52 $3.38
18DeepSeek V4 FlashDeepSeek +2 more hosts 51.8 from $0.110
19Gemini 3.6 FlashGoogle 51.6 $3.00
20Claude Sonnet 4.6Anthropic 48.4 $6.00

LMSYS Arena ELO

Crowd head-to-head preference rating from blind human votes. Typically 1300–1550; higher is better.

Top 20 models by LMSYS Chatbot Arena ELO rating, with blended cost
ModelArena ELOCost $/M
1GPT-5.5 ProOpenAI 1510 $67.50
2Claude Fable 5Anthropic 1507 $20.00
3Claude Opus 4.6Anthropic 1498 $10.00
4Qwen3.8-MaxAlibaba 1497 $3.00
5Claude Opus 4.7Anthropic 1493 $10.00
6Muse Spark 1.1Meta 1488 $2.00
7Gemini 3.1 ProGoogle 1487 $4.50
8Kimi K3Moonshot 1485 $6.00
9Gemini 3.6 FlashGoogle 1485 $3.00
10GPT-5.6 SolOpenAI 1482 $11.25
11GPT-5.4 ProOpenAI 1478 $67.50
12GPT-5.5OpenAI 1477 $11.25
13Gemini 3.5 FlashGoogle 1476 $3.38
14GPT-5.2 ChatOpenAI 1476 $4.81
15Grok 4.20xAI 1475 $1.56
16Gemini 3 Flash PreviewGoogle 1473 $1.12
17Claude Opus 4.8Anthropic 1473 $10.00
18Claude Sonnet 4.6Anthropic 1472 $6.00
19GLM-5.2Z.AI 1469 $2.15
20Grok 4.5xAI 1468 $3.00

These two indices use different scales and can't be added together — a model can rank highly on one and be absent from the other. Composite scores complement, and don't replace, the per-benchmark accuracy tables below.

Best reasoning — GPQA Diamond

PhD-level science reasoning. The hardest general reasoning benchmark.

Top 15 models by GPQA Diamond score, with blended cost and performance per dollar
ModelGPQA scoreCost $/MPerf/$
1Claude Sonnet 5Anthropic 96.2% $4.00 20.4
2GPT-5.6 SolOpenAI 94.6% $11.25 7.4
3GPT-5.4 ProOpenAI 94.4% $67.50
4Gemini 3.1 ProGoogle 94.3% $4.50 16.6
5Claude Opus 4.7Anthropic 94.2% $10.00 8.3
6Claude Mythos 5Anthropic 94.1% $20.00 4.4
7Claude Fable 5Anthropic 94.1% $20.00 4.5
8Claude Opus 4.8Anthropic 93.6% $10.00 8.1
9GPT-5.5OpenAI 93.6% $11.25 6.7
10Kimi K3Moonshot 93.5% $6.00 13.9
11MiniMax-M3MiniMax +2 more hosts 93% from $0.520 120.2
12GPT-5.6 TerraOpenAI 92.9% $4.50 19.5
13Qwen3.8-MaxAlibaba 92.6% $3.00
14GPT-5.2OpenAI 92.4% $4.81 15.8
15GPT-5.6 LunaOpenAI 92.3% $0.450 201.6

Best coding — SWE-Bench

Real-world software engineering tasks — resolving GitHub issues across popular repos.

Top 15 models by SWE-Bench score, with blended cost and performance per dollar
ModelSWE-BenchCost $/MPerf/$
1Claude Opus 5Anthropic 97% $10.00 8.4
2GPT-5.6 SolOpenAI 96.2% $11.25 7.4
3Claude Mythos 5Anthropic 95.5% $20.00 4.4
4Claude Fable 5Anthropic 95% $20.00 4.5
5Kimi K3Moonshot 93.4% $6.00 13.9
6GPT-5.6 LunaOpenAI 93% $0.450 201.6
7Claude Opus 4.8Anthropic 88.6% $10.00 8.1
8Claude Opus 4.7Anthropic 87.6% $10.00 8.3
9Claude Sonnet 5Anthropic 85.2% $4.00 20.4
10GPT-5.3 CodexOpenAI 85% $4.81 14.8
11Claude Sonnet 4.5Anthropic 82% $6.00 12.5
12DeepSeek V4 ProDeepSeek +2 more hosts 80.6% from $0.540 135.2
13GPT-5.2OpenAI 80% $4.81 15.8
14DeepSeek V4 FlashDeepSeek +2 more hosts 79% from $0.110 442.9
15Gemini 3 Flash PreviewGoogle 78% $1.12

Terminal-Bench 2.1

Multi-step tasks in a terminal environment.

Top 15 models by Terminal-Bench 2.1 score, with blended cost and performance per dollar
ModelTerminal-Bench 2.1Cost $/MPerf/$
1GPT-5.6 SolOpenAI 88.8% $11.25 7.4
2Kimi K3Moonshot 88.3% $6.00 13.9
3Claude Mythos 5Anthropic 88% $20.00 4.4
4Claude Fable 5Anthropic 84.3% $20.00 4.5
5GPT-5.6 LunaOpenAI 84.3% $0.450 201.6
6Grok 4.5xAI 83.3% $3.00
7GPT-5.5OpenAI 82.7% $11.25 6.7
8GPT-5.6 TerraOpenAI 82.5% $4.50 19.5
9GLM-5.2Z.AI +2 more hosts 81% from $2.15 35.3
10Claude Sonnet 5Anthropic 80.4% $4.00 20.4
11GPT-5.3 CodexOpenAI 77.3% $4.81 14.8
12Gemini 3.5 FlashGoogle 76.2% $3.38 20.8
13Claude Opus 4.8Anthropic 74.6% $10.00 8.1
14Gemini 3.1 ProGoogle 70.3% $4.50 16.6
15Claude Opus 4.7Anthropic 69.4% $10.00 8.3

LiveCodeBench

Contamination-resistant code generation.

Top 14 models by LiveCodeBench score, with blended cost and performance per dollar
ModelLiveCodeBenchCost $/MPerf/$
1DeepSeek V4 ProDeepSeek +2 more hosts 93.5% from $0.540 135.2
2DeepSeek V4 FlashDeepSeek +2 more hosts 91.6% from $0.110 442.9
3Kimi K2.5Moonshot +1 more host 85% from $1.20 52.4
4Grok 4xAI 79% $6.00 10.5
5Claude Opus 4.6Anthropic 76% $10.00 7.7
6o3-miniOpenAI 74.1% $1.93 34.2
7Claude Sonnet 4.6Anthropic 72.4% $6.00 11.7
8Gemini 2.5 ProGoogle 69% $3.44 17.8
9GPT-OSS 20BFireworks +1 more host 69% from $0.130 490.2
10GPT-OSS 120BFireworks +2 more hosts 69% from $0.260 249.5
11Gemini 2.5 FlashGoogle 63.5% $0.850 59.1
12GPT-4.1OpenAI 52% $3.50 16.3
13Llama 4 MaverickDeepInfra 41% $0.350 186
14Llama 4 ScoutMeta +2 more hosts 32.8% from $0.150 279.4

MCP Atlas

Tool use over Model Context Protocol servers.

Top 15 models by MCP Atlas score, with blended cost and performance per dollar
ModelMCP AtlasCost $/MPerf/$
1Kimi K3Moonshot 84.2% $6.00 13.9
2Gemini 3.5 FlashGoogle 83.6% $3.38 20.8
3Claude Fable 5Anthropic 83.3% $20.00 4.5
4Claude Opus 4.8Anthropic 82.2% $10.00 8.1
5Claude Opus 4.7Anthropic 79.1% $10.00 8.3
6GLM-5.2Z.AI +2 more hosts 77% from $2.15 35.3
7Claude Opus 4.6Anthropic 76.8% $10.00 7.7
8GPT-5.5OpenAI 75.3% $11.25 6.7
9MiniMax-M3MiniMax +2 more hosts 74.2% from $0.520 120.2
10DeepSeek V4 ProDeepSeek +2 more hosts 73.6% from $0.540 135.2
11Claude Opus 4.5Anthropic 69.8% $10.00 7.1
12Claude Sonnet 4.6Anthropic 69.5% $6.00 11.7
13Gemini 3.1 ProGoogle 69.2% $4.50 16.6
14DeepSeek V4 FlashDeepSeek +2 more hosts 69% from $0.110 442.9
15GPT-5.2OpenAI 67.6% $4.81 15.8

BrowseComp

Agentic web search and information extraction.

Top 12 models by BrowseComp score, with blended cost and performance per dollar
ModelBrowseCompCost $/MPerf/$
1GPT-5.6 SolOpenAI 92.2% $11.25 7.4
2Kimi K3Moonshot 91.2% $6.00 13.9
3Claude Opus 5Anthropic 90.8% $10.00 8.4
4Claude Fable 5Anthropic 88% $20.00 4.5
5Gemini 3.1 ProGoogle 85.9% $4.50 16.6
6DeepSeek V4 FlashDeepSeek +2 more hosts 85.9% from $0.110 442.9
7Claude Sonnet 5Anthropic 84.7% $4.00 20.4
8GPT-5.5OpenAI 84.4% $11.25 6.7
9MiniMax-M3MiniMax +2 more hosts 83.5% from $0.520 120.2
10DeepSeek V4 ProDeepSeek +2 more hosts 83.4% from $0.540 135.2
11Kimi K2.6Moonshot +1 more host 83.2% from $1.71 44.5
12Claude Sonnet 4.6Anthropic 76.2% $6.00 11.7

OSWorld-Verified

Computer use — real GUI tasks on a desktop OS.

Top 12 models by OSWorld-Verified score, with blended cost and performance per dollar
ModelOSWorld-VerifiedCost $/MPerf/$
1Claude Fable 5Anthropic 85% $20.00 4.5
2Claude Opus 4.8Anthropic 83.4% $10.00 8.1
3Claude Sonnet 5Anthropic 81.2% $4.00 20.4
4GPT-5.5OpenAI 78.7% $11.25 6.7
5Claude Sonnet 4.6Anthropic 78.5% $6.00 11.7
6Gemini 3.5 FlashGoogle 78.4% $3.38 20.8
7Claude Opus 4.7Anthropic 78% $10.00 8.3
8Gemini 3.1 ProGoogle 76.2% $4.50 16.6
9Kimi K2.6Moonshot +1 more host 73.1% from $1.71 44.5
10MiniMax-M3MiniMax +2 more hosts 70.1% from $0.520 120.2
11GPT-5.3 CodexOpenAI 64.7% $4.81 14.8
12GPT-5.6 SolOpenAI 62.6% $11.25 7.4

Best math — AIME 2025

American Invitational Mathematics Examination problems — competition-level math.

Top 15 models by AIME 2025 score, with blended cost and performance per dollar
ModelAIME scoreCost $/MPerf/$
1GPT-5.2OpenAI 100% $4.81 15.8
2GPT-5.2 ProOpenAI 100% $57.75
3Claude Opus 4.6Anthropic 99.8% $10.00 7.7
4GPT-OSS 20BFireworks +1 more host 98.7% from $0.130 490.2
5GPT-OSS 120BFireworks +2 more hosts 97.9% from $0.260 249.5
6Claude Haiku 4.5Anthropic 96.3% $2.00 32.8
7Kimi K2.5Moonshot +1 more host 96.1% from $1.20 52.4
8GPT-5.4OpenAI 95.5% $5.62 14
9o4-miniOpenAI 92.7% $1.93 35.9
10Grok 4xAI 91.7% $6.00 10.5
11Gemini 2.5 ProGoogle 88% $3.44 17.8
12Claude Sonnet 4.5Anthropic 87% $6.00 12.5
13o3-miniOpenAI 86.5% $1.93 34.2
14Grok 4.3xAI 85% $1.56 46.3
15Claude Sonnet 4.6Anthropic 83% $6.00 11.7

Fastest endpoints — inference speed

Tokens per second on hosted inference — the one board here ranked by endpoint rather than by model, because speed is a property of the host's serving stack, not of the weights. The same model can appear more than once at different hosts, and the difference between those rows is the point. Speed matters for real-time applications and high-throughput pipelines.

Top 15 hosted endpoints by inference speed in tokens per second, with blended cost — ranked per host, so one model may appear more than once
ModelTokens/secCost $/M
1Llama 4 ScoutMeta 2600 t/s $0.170
2Llama 4 ScoutGroq 2600 t/s $0.170
3Llama 4 ScoutDeepInfra 2600 t/s $0.150
4Llama 3.3 70BMeta 2500 t/s $0.640
5Llama 3.3 70B VersatileGroq 2500 t/s $0.640
6Llama 3.3 70BTogether 2500 t/s $1.04
7GPT-OSS 120BFireworks 1828.8 t/s $0.260
8Llama 3.1 8BMeta 1800 t/s $0.060
9Llama 3.1 8B InstantGroq 1800 t/s $0.060
10GPT-OSS 20BFireworks 941.5 t/s $0.130
11Gemini 3.1 Flash-LiteGoogle 600 t/s $0.560
12GPT-OSS 20BGroq 564 t/s $0.130
13Qwen 3.6 27BGroq 472.4 t/s $1.20
14Gemini 2.5 Flash-LiteGoogle 450 t/s $0.180
15Granite 4 H SmallIBM 414.5 t/s $0.110

Hardest benchmark — Humanity's Last Exam

Expert-level questions across 100+ fields. No model scores above 65% — the frontier of AI capability.

Top 10 models by Humanity's Last Exam score, with blended cost
ModelHLE scoreCost $/M
1Claude Opus 5Anthropic 64.7% $10.00
2Claude Mythos 5Anthropic 64.5% $20.00
3Claude Opus 4.8Anthropic 57.9% $10.00
4Claude Sonnet 5Anthropic 57.4% $4.00
5Kimi K3Moonshot 56% $6.00
6GLM-5.2Z.AI +2 more hosts 54.7% from $2.15
7Kimi K2.6Moonshot +1 more host 54% from $1.71
8DeepSeek V4 FlashDeepSeek +2 more hosts 51.6% from $0.110
9DeepSeek V4 ProDeepSeek +2 more hosts 48.2% from $0.540
10GPT-5.6 SolOpenAI 47.2% $11.25

Methodology

how the scores and rankings are computed

Benchmark sources

Vellum LLM Leaderboard, Artificial Analysis, official model technical reports, and HuggingFace Open LLM Leaderboard. Scores are the latest publicly reported results as of August 2026, refreshed as new evaluations are published.

Performance per dollar

Shown two ways. The raw Perf/$ column is avg_benchmark_score / blended_cost_per_mtok (blended cost = (3×input + 1×output) ÷ 4 per million tokens). The Best Value ranking is driven instead by the comparability-normalized Norm Value: each model's benchmark strength is first converted to a percentile rank against the whole tracked population (avg_method = percentile_rank_vs_population), then divided by the same blended cost — so a model tested only on easy benchmarks can't post an inflated ratio.

A Norm Value is a percentile-derived value figure, not a % accuracy — only the raw Avg Score and per-benchmark columns are accuracy percentages.

Comparability rule

A model's "average" only covers the benchmarks it has actually been tested on, and the tracked benchmarks differ wildly in difficulty (nobody scores above ~65% on Humanity's Last Exam; AIME scores run into the high 90s), so averaging over different subsets isn't apples-to-apples. Every avg-based ranking and the scatter above therefore require at least 3 filled benchmarks and display the benchmark count (n). Single-benchmark tables (GPQA, SWE-Bench, Terminal-Bench, AIME, HLE and the rest) are unaffected — one shared benchmark is directly comparable.

One row per model

Open-weight models are often sold by several hosts at once — GLM-5.2 by Z.AI, Fireworks and Together. A benchmark score belongs to the weights, so all three score identically, and listing them separately would spend three ranking slots on one answer. Every board here shows a model once, at its cheapest host, with the remaining hosts noted on the row; the price reads “from” whenever more than one host sells it. The counts above follow the same rule — a model sold by three hosts counts once.

The one exception is Fastest endpoints. Tokens per second is a property of the host's serving stack rather than of the weights — Z.AI runs GLM-5.2 at a different speed from Fireworks — so that board ranks endpoints, and a model can legitimately appear on it more than once.

What each benchmark measures

GPQA Diamond
PhD-level questions in biology, chemistry, and physics
AIME 2025
American Invitational Mathematics Examination (competition math)
SWE-Bench
Resolving real GitHub issues in popular Python repositories
Humanity's Last Exam
Expert-level questions across 100+ academic fields
ARC-AGI 2
Visual reasoning and abstract pattern completion
MMMLU
Multilingual version of MMLU (57 subjects, multiple languages)
HumanEval
Python code generation from natural language descriptions
MATH 500
Mathematical problem solving across 5 difficulty levels
BFCL
Berkeley Function Calling Leaderboard (tool use & function calling)

Access this data programmatically via the benchmarks API endpoint or the API documentation.