ModelPriceWatch.com
Last scan 2026-09-22 Models tracked 267 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

LLM performance rankings

Benchmark scores and performance-per-dollar across 110 benchmarked models. Updated with the latest public evaluations — every ranking pairs the score with what the model actually costs to run.

Models benchmarked

110

of 231 tracked models — scores from public leaderboards. One model counts once, however many hosts sell it.

Benchmarks tracked

16

GPQA Diamond · SWE-Bench Verified · Terminal-Bench 2.1 · LiveCodeBench · MCP Atlas · BrowseComp · OSWorld-Verified · Humanity's Last Exam · ARC-AGI 2 · AIME 2025 · MMMLU · BFCL · HumanEval · MATH 500 · AutoBench · AIME 2026

Providers covered

26

each score keeps its source — how we count

Cost vs quality — the value frontier

every dot is a model page — hover for the numbers

50 models plotted by blended cost (log scale) against average benchmark score — hover any dot for details, click to open the model. The dashed line is the value frontier: no model is both cheaper and better than a point on it. Cheapest model scoring 80+: GPT-5.6 Luna at $0.450/Mtok. Only models with ≥3 of the 9 tracked benchmarks are averaged (21 thinly-tested models excluded — a one-benchmark "average" isn't comparable).

30 40 50 60 70 80 90 $0.05 $0.1 $0.3 $1 $3 $10 Blended cost $/Mtok (log scale) → Avg benchmark score → DeepSeek V4 Flash GPT-5.6 Luna Llama 3.1 8B
Closed weights Open weights Best value Value frontier MODELPRICEWATCH.COM · 2026-09-22

Best value — performance per dollar

norm value drives the ranking — top 25 shown

Ranked by comparability-normalized value — each model's benchmark strength is scored as a percentile rank against the whole tracked population, then divided by blended cost ($/Mtok), so a model tested only on easy benchmarks can't buy its way to the top. The Perf/$ column shows the raw ratio for reference; the Norm value column is what drives the ranking. Ranked models have ≥3 of the 9 tracked benchmarks — the "n" column shows how many; averages over one or two benchmarks aren't comparable and are excluded.

Best-value LLM ranking — benchmark strength per blended dollar, top 25 of 110 benchmarked models
Model Norm value Avg score (n) Blended $/M Perf/$ Value
1DeepSeek V4 FlashDeepInfra +2 more hosts 515.5 77.5% (9) from $0.110 688.9
2GLM-5.3-FlashZ.AI 343.8 69.8% (6) $0.240 293.9
3Solar Pro 4Upstage 265.6 57% (3) $0.160 361.9
4GPT-OSS 20BFireworks +1 more host 223.8 62.5% (6) from $0.130 490.2
5GPT-5.6 LunaOpenAI 150.2 90.7% (8) $0.450 201.6
6GPT-OSS 120BTogether +2 more hosts 136.9 65.5% (6) from $0.260 249.5
7Llama 4 ScoutDeepInfra +2 more hosts 128 46.8% (12) from $0.150 312
8DeepSeek V4.1 FlashDeepSeek 119 56.8% (5) $0.520 108.2
9MiniMax-M3MiniMax +2 more hosts 83.8 63.1% (17) from $0.520 120.2
10Llama 4 MaverickDeepInfra 79.4 65.1% (5) $0.350 186
11Llama 3.1 8BMeta 70 34% (9) $0.060 591.3
12Llama 3.1 8B InstantGroq 70 34% (9) $0.060 591.3
13Gemini 3.7 FlashGoogle 53.1 85.8% (5) $1.50 57.2
14Muse Spark 1.3Meta 44.6 67.5% (6) $2.00 33.8
15Mistral Large 3Mistral 44.5 53.6% (12) $0.750 71.5
16Grok 4.3xAI 41.8 72.4% (12) $1.56 46.3
17Muse Glimmer 30BFireworks +1 more host 41.2 37% (6) from $0.640 58
18Qwen3.8-27BAlibaba 40.8 34% (5) $1.12 30.2
19GPT-4.1 nanoOpenAI 39.4 58.6% (3) $0.180 334.9
20Muse Spark 1.2Meta 38.8 62% (6) $2.00 31
21GLM-5.3Z.AI +2 more hosts 37.7 42% (5) from $2.15 19.5
22Claude Haiku 3.5Anthropic 36.7 78.8% (3) $1.60 49.2
23DeepSeek V4 ProDeepInfra +3 more hosts 36.2 73.5% (14) from $1.62 45.2
24Gemini 2.5 FlashGoogle 35.3 50.2% (10) $0.850 59.1
25Kimi K2.6Fireworks +1 more host 35.1 76.2% (7) from $1.71 44.5
Norm value = percentile-rank benchmark strength ÷ blended $/Mtok. How we rank MODELPRICEWATCH.COM · 2026-09-22

Composite intelligence scores

published values only — two scales, never added together

Two independent, cross-model composite indices. Unlike the per-benchmark tables below, these summarise overall capability in a single number — and for the newest family flagships (which the per-benchmark leaderboards don't yet cover) they are often the only real, independent scores that exist. We show published values only; a means no composite has been released for that model. Sources: Artificial Analysis Intelligence Index and LMSYS Chatbot Arena ELO (UC Berkeley).

Artificial Analysis Intelligence Index

A 0–100 blended index across reasoning, math, coding and knowledge evals. Higher is better.

Top 20 models by Artificial Analysis Intelligence Index, with blended cost
ModelAA indexCost $/M
1Claude Fable 5.1Anthropic 53.4 $20.00
2GPT-6 AstraOpenAI 52.8 $20.00
3DeepSeek V4 Flash Vision ExpDeepSeek 51 $0.660
4Claude Opus 5Anthropic 50.7 $10.00
5Claude Fable 5Anthropic 49.7 $20.00
6Muse Spark 1.3Meta 48.2 $2.00
7GPT-5.6 SolOpenAI 47.1 $8.00
8Kimi K2.6Fireworks 45.1 $1.71
9GLM-5.3Z.AI +2 more hosts 44.9 from $2.15
10Grok 4.6xAI 44.4 $3.00
11GPT-5.3 CodexOpenAI 44.3 $4.81
12Kimi K3Moonshot 43.8 $6.00
13GPT-5.6 TerraOpenAI 42.3 $4.50
14GPT-5.2OpenAI 42.2 $4.81
15Claude Opus 4.8Anthropic 42 $10.00
16Solar Pro 4Upstage 42 $0.160
17GLM-5.3-FlashZ.AI 41.9 $0.240
18Gemini 3.8 FlashGoogle 41.2 $1.50
19Qwen3.6-PlusFireworks 40.5 $1.12
20Qwen3.8-MaxAlibaba 40.3 $3.00

LMSYS Arena ELO

Crowd head-to-head preference rating from blind human votes. Typically 1300–1550; higher is better.

Top 20 models by LMSYS Chatbot Arena ELO rating, with blended cost
ModelArena ELOCost $/M
1GPT-5.5 ProOpenAI 1510 $67.50
2Claude Fable 5Anthropic 1506 $20.00
3Muse Spark 1.2Meta 1500 $2.00
4Claude Fable 5.1Anthropic 1499 $20.00
5Claude Opus 4.6Anthropic 1497 $10.00
6Claude Opus 4.7Anthropic 1494 $10.00
7Muse Spark 1.3Meta 1493 $2.00
8Gemini 3.8 FlashGoogle 1493 $1.50
9Claude Opus 5Anthropic 1493 $10.00
10Muse Spark 1.1Meta 1493 $2.00
11Gemini 3.7 FlashGoogle 1490 $1.50
12Gemini 3.1 ProGoogle 1487 $4.50
13Kimi K3Moonshot 1485 $6.00
14GPT-5.6 SolOpenAI 1484 $8.00
15GLM-5.3Z.AI 1483 $2.15
16Claude Opus 4.8Anthropic 1481 $10.00
17Qwen3.8-MaxAlibaba 1481 $3.00
18Gemini 3.6 FlashGoogle 1480 $1.50
19GPT-6 AstraOpenAI 1480 $20.00
20GPT-5.4 ProOpenAI 1478 $67.50

These two indices use different scales and can't be added together — a model can rank highly on one and be absent from the other. Composite scores complement, and don't replace, the per-benchmark accuracy tables below.

The per-benchmark boards below rank independently measured scores. A maker's own number is flagged "vendor", listed after the independent rows, and never competes for a rank.

Best reasoning — GPQA Diamond

PhD-level science reasoning. The hardest general reasoning benchmark.

Top 15 models by GPQA Diamond score, with blended cost and performance per dollar
ModelGPQA scoreCost $/MPerf/$
1Claude Sonnet 5Anthropic 96.2% $4.00 20.4
2GPT-6 AstraOpenAI 96% $20.00 4.2
3GPT-5.6 SolOpenAI 94.6% $8.00 10.5
4Gemini 3.1 ProGoogle 94.3% $4.50 16.6
5Claude Opus 4.7Anthropic 94.2% $10.00 8.3
6Claude Mythos 5Anthropic 94.1% $20.00 4.4
7Claude Fable 5Anthropic 94.1% $20.00 4.5
8Claude Opus 4.8Anthropic 93.6% $10.00 8.1
9GPT-5.5OpenAI 93.6% $11.25 6.7
10Kimi K3Moonshot 93.5% $6.00 13.9
11Claude Fable 5.1Anthropic 93.4% $20.00 4.2
12MiniMax-M3MiniMax +2 more hosts 93% from $0.520 120.2
13GPT-5.6 TerraOpenAI 92.9% $4.50 19.5
14GPT-5.2OpenAI 92.4% $4.81 15.8
15GPT-5.6 LunaOpenAI 92.3% $0.450 201.6

Best coding — SWE-Bench

Real-world software engineering tasks — resolving GitHub issues across popular repos.

Top 15 models by SWE-Bench score, with blended cost and performance per dollar
ModelSWE-BenchCost $/MPerf/$
1Claude Opus 5Anthropic 97% $10.00 8.4
2GPT-5.6 SolOpenAI 96.2% $8.00 10.5
3Claude Mythos 5Anthropic 95.5% $20.00 4.4
4Claude Fable 5Anthropic 95% $20.00 4.5
5Kimi K3Moonshot 93.4% $6.00 13.9
6GPT-5.6 LunaOpenAI 93% $0.450 201.6
7Claude Opus 4.8Anthropic 88.6% $10.00 8.1
8Claude Opus 4.7Anthropic 87.6% $10.00 8.3
9Claude Sonnet 5Anthropic 85.2% $4.00 20.4
10Claude Sonnet 4.5Anthropic 82% $6.00 12.5
11GPT-5.4OpenAI 75% $5.62 14
12o4-miniOpenAI 68% $1.93 35.9
13Grok 4.3xAI 65% $1.56 46.3
14Grok 4xAI 55% $6.00 10.5
15Gemini 2.5 ProGoogle 50% $3.44 17.8

Terminal-Bench 2.1

Multi-step tasks in a terminal environment.

Top 15 models by Terminal-Bench 2.1 score, with blended cost and performance per dollar
ModelTerminal-Bench 2.1Cost $/MPerf/$
1Claude Fable 5.1Anthropic 91.4% $20.00 4.2
2GPT-5.6 SolOpenAI 88.8% $8.00 10.5
3Kimi K3Moonshot 88.3% $6.00 13.9
4Claude Mythos 5Anthropic 88% $20.00 4.4
5Muse Spark 1.3Meta 86% $2.00 33.8
6Gemini 3.7 FlashGoogle 85.8% $1.50 57.2
7Claude Fable 5Anthropic 84.3% $20.00 4.5
8GPT-5.6 LunaOpenAI 84.3% $0.450 201.6
9GLM-5.3-FlashZ.AI 84.3% $0.240 293.9
10GPT-5.5OpenAI 82.7% $11.25 6.7
11GPT-5.6 TerraOpenAI 82.5% $4.50 19.5
12GLM-5.2Z.AI +2 more hosts 81% from $2.15 35.3
13Claude Sonnet 5Anthropic 80.4% $4.00 20.4
14Muse Spark 1.2Meta 80% $2.00 31
15Muse Spark 1.1Meta 78% $2.00 30.8

LiveCodeBench

Contamination-resistant code generation.

Top 15 models by LiveCodeBench score, with blended cost and performance per dollar
ModelLiveCodeBenchCost $/MPerf/$
1DeepSeek V4 ProDeepSeek +3 more hosts 93.5% from $1.62 37.1
2DeepSeek V4 FlashDeepSeek +2 more hosts 91.6% from $0.110 117.4
3Claude Fable 5.1Anthropic 90.5% $20.00 4.2
4Kimi K2.5Moonshot +1 more host 85% from $1.20 52.4
5Grok 4xAI 79% $6.00 10.5
6Claude Opus 4.6Anthropic 76% $10.00 7.7
7o3-miniOpenAI 74.1% $1.93 34.2
8Claude Sonnet 4.6Anthropic 72.4% $6.00 11.7
9Gemini 2.5 ProGoogle 69% $3.44 17.8
10GPT-OSS 20BFireworks +1 more host 69% from $0.130 490.2
11GPT-OSS 120BFireworks +2 more hosts 69% from $0.260 249.5
12Gemini 2.5 FlashGoogle 63.5% $0.850 59.1
13GPT-4.1OpenAI 52% $3.50 16.3
14Llama 4 MaverickDeepInfra 41% $0.350 186
15Llama 4 ScoutMeta +2 more hosts 32.8% from $0.150 279.4

MCP Atlas

Tool use over Model Context Protocol servers.

Top 15 models by MCP Atlas score, with blended cost and performance per dollar
ModelMCP AtlasCost $/MPerf/$
1Kimi K3Moonshot 84.2% $6.00 13.9
2Gemini 3.5 FlashGoogle 83.6% $3.38 20.8
3Claude Fable 5Anthropic 83.3% $20.00 4.5
4Claude Opus 4.8Anthropic 82.2% $10.00 8.1
5Claude Opus 4.7Anthropic 79.1% $10.00 8.3
6GLM-5.2Z.AI +2 more hosts 77% from $2.15 35.3
7Claude Opus 4.6Anthropic 76.8% $10.00 7.7
8GPT-5.5OpenAI 75.3% $11.25 6.7
9MiniMax-M3MiniMax +2 more hosts 74.2% from $0.520 120.2
10DeepSeek V4 ProDeepSeek +3 more hosts 73.6% from $1.62 37.1
11Claude Opus 4.5Anthropic 69.8% $10.00 7.1
12Claude Sonnet 4.6Anthropic 69.5% $6.00 11.7
13Gemini 3.1 ProGoogle 69.2% $4.50 16.6
14DeepSeek V4 FlashDeepSeek +2 more hosts 69% from $0.110 117.4
15GPT-5.2OpenAI 67.6% $4.81 15.8

BrowseComp

Agentic web search and information extraction.

Top 15 models by BrowseComp score, with blended cost and performance per dollar
ModelBrowseCompCost $/MPerf/$
1GPT-5.6 SolOpenAI 92.2% $8.00 10.5
2GPT-6 AstraOpenAI 91.5% $20.00 4.2
3Kimi K3Moonshot 91.2% $6.00 13.9
4Claude Opus 5Anthropic 90.8% $10.00 8.4
5Claude Fable 5Anthropic 88% $20.00 4.5
6Gemini 3.1 ProGoogle 85.9% $4.50 16.6
7DeepSeek V4 FlashDeepSeek +2 more hosts 85.9% from $0.110 117.4
8Claude Sonnet 5Anthropic 84.7% $4.00 20.4
9GPT-5.5OpenAI 84.4% $11.25 6.7
10MiniMax-M3MiniMax +2 more hosts 83.5% from $0.520 120.2
11DeepSeek V4 ProDeepSeek +3 more hosts 83.4% from $1.62 37.1
12Kimi K2.6Moonshot +1 more host 83.2% from $1.71 44.5
13Claude Sonnet 4.6Anthropic 76.2% $6.00 11.7
14Inkling vendorThinking Machines 77.1% $1.76
15Solar Pro 4 vendorUpstage 49.2% $0.160 361.9

OSWorld-Verified

Computer use — real GUI tasks on a desktop OS.

Top 15 models by OSWorld-Verified score, with blended cost and performance per dollar
ModelOSWorld-VerifiedCost $/MPerf/$
1Claude Fable 5Anthropic 85% $20.00 4.5
2Claude Opus 4.8Anthropic 83.4% $10.00 8.1
3Claude Sonnet 5Anthropic 81.2% $4.00 20.4
4GPT-5.5OpenAI 78.7% $11.25 6.7
5Claude Sonnet 4.6Anthropic 78.5% $6.00 11.7
6Gemini 3.5 FlashGoogle 78.4% $3.38 20.8
7Claude Opus 4.7Anthropic 78% $10.00 8.3
8Gemini 3.1 ProGoogle 76.2% $4.50 16.6
9Kimi K2.6Moonshot +1 more host 73.1% from $1.71 44.5
10MiniMax-M3MiniMax +2 more hosts 70.1% from $0.520 120.2
11GPT-5.3 CodexOpenAI 64.7% $4.81 14.8
12GPT-5.6 SolOpenAI 62.6% $8.00 10.5
13Qwen3.8-27B vendorAlibaba 84.3% $1.12 30.2
14Gemini 3.6 Flash vendorGoogle 83% $1.50
15Muse Spark 1.1 vendorMeta 80.8% $2.00 30.8

Best math — AIME 2025

American Invitational Mathematics Examination problems — competition-level math.

Top 15 models by AIME 2025 score, with blended cost and performance per dollar
ModelAIME scoreCost $/MPerf/$
1GPT-5.2OpenAI 100% $4.81 15.8
2Claude Opus 4.6Anthropic 99.8% $10.00 7.7
3GPT-OSS 20BFireworks +1 more host 98.7% from $0.130 490.2
4GPT-OSS 120BFireworks +2 more hosts 97.9% from $0.260 249.5
5Claude Haiku 4.5Anthropic 96.3% $2.00 32.8
6Kimi K2.5Moonshot +1 more host 96.1% from $1.20 52.4
7GPT-5.4OpenAI 95.5% $5.62 14
8o4-miniOpenAI 92.7% $1.93 35.9
9Grok 4xAI 91.7% $6.00 10.5
10Gemini 2.5 ProGoogle 88% $3.44 17.8
11Claude Sonnet 4.5Anthropic 87% $6.00 12.5
12o3-miniOpenAI 86.5% $1.93 34.2
13Grok 4.3xAI 85% $1.56 46.3
14Claude Sonnet 4.6Anthropic 83% $6.00 11.7
15Claude Opus 4.1Anthropic 78% $30.00 2.1

Fastest endpoints — inference speed

Tokens per second on hosted inference — the one board here ranked by endpoint rather than by model, because speed is a property of the host's serving stack, not of the weights. The same model can appear more than once at different hosts, and the difference between those rows is the point. Speed matters for real-time applications and high-throughput pipelines.

Top 15 hosted endpoints by inference speed in tokens per second, with blended cost — ranked per host, so one model may appear more than once
ModelTokens/secCost $/M
1Llama 4 ScoutMeta 2600 t/s $0.170
2Llama 4 ScoutGroq 2600 t/s $0.170
3Llama 4 ScoutDeepInfra 2600 t/s $0.150
4Llama 3.3 70BMeta 2500 t/s $0.640
5Llama 3.3 70B VersatileGroq 2500 t/s $0.640
6Llama 3.3 70BTogether 2500 t/s $1.04
7GPT-OSS 120BFireworks 1828.8 t/s $0.260
8Llama 3.1 8BMeta 1800 t/s $0.060
9Llama 3.1 8B InstantGroq 1800 t/s $0.060
10GPT-OSS 20BFireworks 941.5 t/s $0.130
11Gemini 3.1 Flash-LiteGoogle 600 t/s $0.560
12GPT-OSS 20BGroq 564 t/s $0.130
13Qwen3.6-27BGroq 472.4 t/s $1.20
14Gemini 2.5 Flash-LiteGoogle 450 t/s $0.180
15Granite 4 H SmallIBM 414.5 t/s $0.110

Hardest benchmark — Humanity's Last Exam

Expert-level questions across 100+ fields. No model scores above 65% — the frontier of AI capability.

Top 10 models by Humanity's Last Exam score, with blended cost
ModelHLE scoreCost $/M
1Claude Opus 5Anthropic 64.7% $10.00
2Claude Mythos 5Anthropic 64.5% $20.00
3Claude Fable 5.1Anthropic 59.1% $20.00
4Claude Opus 4.8Anthropic 57.9% $10.00
5Claude Sonnet 5Anthropic 57.4% $4.00
6GPT-6 AstraOpenAI 57.2% $20.00
7Kimi K3Moonshot 56% $6.00
8GLM-5.3-FlashZ.AI 55.3% $0.240
9GLM-5.2Z.AI +2 more hosts 54.7% from $2.15
10Kimi K2.6Moonshot +1 more host 54% from $1.71

Methodology

how the scores and rankings are computed

Benchmark sources

Vellum LLM Leaderboard, Artificial Analysis, official model technical reports, and HuggingFace Open LLM Leaderboard. Scores are the latest publicly reported results as of September 2026, refreshed as new evaluations are published.

Performance per dollar

Shown two ways. The raw Perf/$ column is avg_benchmark_score / blended_cost_per_mtok (blended cost = (3×input + 1×output) ÷ 4 per million tokens). The Best Value ranking is driven instead by the comparability-normalized Norm Value: each model's benchmark strength is first converted to a percentile rank against the whole tracked population (avg_method = percentile_rank_vs_population), then divided by the same blended cost — so a model tested only on easy benchmarks can't post an inflated ratio.

A Norm Value is a percentile-derived value figure, not a % accuracy — only the raw Avg Score and per-benchmark columns are accuracy percentages.

Comparability rule

A model's "average" only covers the benchmarks it has actually been tested on, and the tracked benchmarks differ wildly in difficulty (nobody scores above ~65% on Humanity's Last Exam; AIME scores run into the high 90s), so averaging over different subsets isn't apples-to-apples. Every avg-based ranking and the scatter above therefore require at least 3 filled benchmarks and display the benchmark count (n). Single-benchmark tables (GPQA, SWE-Bench, Terminal-Bench, AIME, HLE and the rest) are unaffected — one shared benchmark is directly comparable.

One row per model

Open-weight models are often sold by several hosts at once — GLM-5.2 by Z.AI, Fireworks and Together. A benchmark score belongs to the weights, so all three score identically, and listing them separately would spend three ranking slots on one answer. Every board here shows a model once, at its cheapest host, with the remaining hosts noted on the row; the price reads “from” whenever more than one host sells it. The counts above follow the same rule — a model sold by three hosts counts once.

The one exception is Fastest endpoints. Tokens per second is a property of the host's serving stack rather than of the weights — Z.AI runs GLM-5.2 at a different speed from Fireworks — so that board ranks endpoints, and a model can legitimately appear on it more than once.

What each benchmark measures

GPQA Diamond
PhD-level questions in biology, chemistry, and physics
AIME 2025
American Invitational Mathematics Examination (competition math)
SWE-Bench
Resolving real GitHub issues in popular Python repositories
Humanity's Last Exam
Expert-level questions across 100+ academic fields
ARC-AGI 2
Visual reasoning and abstract pattern completion
MMMLU
Multilingual version of MMLU (57 subjects, multiple languages)
HumanEval
Python code generation from natural language descriptions
MATH 500
Mathematical problem solving across 5 difficulty levels
BFCL
Berkeley Function Calling Leaderboard (tool use & function calling)

Access this data programmatically via the benchmarks API endpoint or the API documentation.