LLM performance rankings
Benchmark scores and performance-per-dollar across 89 benchmarked models. Updated with the latest public evaluations — every ranking pairs the score with what the model actually costs to run.
Models benchmarked
89
of 166 tracked models — scores from public leaderboards. One model counts once, however many hosts sell it.
Benchmarks tracked
9
GPQA · AIME · SWE-Bench · HLE · ARC-AGI 2 · MMMLU · HumanEval · MATH 500 · BFCL
Cost vs quality — the value frontier
every dot is a model page — hover for the numbers37 models plotted by blended cost (log scale) against average benchmark score — hover any dot for details, click to open the model. The dashed line is the value frontier: no model is both cheaper and better than a point on it. Cheapest model scoring 80+: GPT-5.6 Luna at $0.450/Mtok. Only models with ≥3 of the 9 tracked benchmarks are averaged (19 thinly-tested models excluded — a one-benchmark "average" isn't comparable).
Best value — performance per dollar
norm value drives the ranking — top 25 shownRanked by comparability-normalized value — each model's benchmark strength is scored as a percentile rank against the whole tracked population, then divided by blended cost ($/Mtok), so a model tested only on easy benchmarks can't buy its way to the top. The Perf/$ column shows the raw ratio for reference; the Norm value column is what drives the ranking. Ranked models have ≥3 of the 9 tracked benchmarks — the "n" column shows how many; averages over one or two benchmarks aren't comparable and are excluded.
| Model | Norm value | Avg score (n) | Blended $/M | Perf/$ | Value |
|---|---|---|---|---|---|
| 1DeepSeek V4 FlashDeepInfra +2 more hosts | 568.2 | 77.5% (10) | from $0.110 | 688.9 | |
| 2GPT-OSS 20BFireworks +1 more host | 223.8 | 62.5% (7) | from $0.130 | 490.2 | |
| 3GPT-5.6 LunaOpenAI | 174.2 | 90.7% (8) | $0.450 | 201.6 | |
| 4GPT-OSS 120BTogether +2 more hosts | 138.8 | 65.5% (7) | from $0.260 | 249.5 | |
| 5Llama 4 ScoutDeepInfra +2 more hosts | 126 | 46.8% (13) | from $0.150 | 312 | |
| 6DeepSeek V4 ProDeepSeek +2 more hosts | 108.3 | 73.5% (16) | from $0.540 | 135.2 | |
| 7MiniMax-M3MiniMax +2 more hosts | 93.8 | 63.1% (18) | from $0.520 | 120.2 | |
| 8Llama 3.1 8BMeta | 73.3 | 34% (9) | $0.060 | 591.3 | |
| 9Llama 3.1 8B InstantGroq | 73.3 | 34% (9) | $0.060 | 591.3 | |
| 10Llama 4 MaverickDeepInfra | 71.1 | 65.1% (6) | $0.350 | 186 | |
| 11Mistral Large 3Mistral | 45.7 | 53.6% (13) | $0.750 | 71.5 | |
| 12Grok 4.3xAI | 41.1 | 72.4% (13) | $1.56 | 46.3 | |
| 13GPT-4.1 nanoOpenAI | 37.8 | 58.6% (5) | $0.180 | 334.9 | |
| 14Claude Haiku 3.5Anthropic | 37.4 | 78.8% (3) | $1.60 | 49.2 | |
| 15Gemini 2.5 FlashGoogle | 37.1 | 50.2% (10) | $0.850 | 59.1 | |
| 16Kimi K2.5Moonshot +1 more host | 35.7 | 62.9% (12) | from $1.20 | 52.4 | |
| 17GLM-5.2Z.AI +2 more hosts | 35.1 | 76% (9) | from $2.15 | 35.3 | |
| 18Kimi K2.6Moonshot +1 more host | 34.7 | 76.2% (9) | from $1.71 | 44.5 | |
| 19Llama 3.3 70BMeta +2 more hosts | 34.5 | 46.6% (9) | from $0.640 | 72.8 | |
| 20o4-miniOpenAI | 29.2 | 69.2% (10) | $1.93 | 35.9 | |
| 21GPT-4.1 miniOpenAI | 25 | 55.7% (7) | $0.700 | 79.6 | |
| 22Gemini 3.5 FlashGoogle | 22.9 | 70.1% (9) | $3.38 | 20.8 | |
| 23Claude Sonnet 5Anthropic | 20.4 | 81.5% (11) | $4.00 | 20.4 | |
| 24GPT-5.6 TerraOpenAI | 17.7 | 87.7% (6) | $4.50 | 19.5 | |
| 25Claude Haiku 4.5Anthropic | 17.4 | 65.5% (10) | $2.00 | 32.8 |
Composite intelligence scores
published values only — two scales, never added togetherTwo independent, cross-model composite indices. Unlike the per-benchmark tables below, these summarise overall capability in a single number — and for the newest family flagships (which the per-benchmark leaderboards don't yet cover) they are often the only real, independent scores that exist. We show published values only; a — means no composite has been released for that model. Sources: Artificial Analysis Intelligence Index and LMSYS Chatbot Arena ELO (UC Berkeley).
Artificial Analysis Intelligence Index
A 0–100 blended index across reasoning, math, coding and knowledge evals. Higher is better.
| Model | AA index | Cost $/M |
|---|---|---|
| 1Claude Opus 5Anthropic | 63.1 | $10.00 |
| 2Claude Fable 5Anthropic | 62.1 | $20.00 |
| 3GPT-5.6 SolOpenAI | 60.9 | $11.25 |
| 4Kimi K3Moonshot | 59.7 | $6.00 |
| 5Qwen3.8-MaxAlibaba | 58.1 | $3.00 |
| 6Claude Opus 4.8Anthropic | 57.3 | $10.00 |
| 7Muse Spark 1.2Meta | 56.8 | $2.00 |
| 8GPT-5.6 TerraOpenAI | 56.6 | $4.50 |
| 9GPT-5.5OpenAI | 56.3 | $11.25 |
| 10Grok 4.5xAI | 55.8 | $3.00 |
| 11Claude Sonnet 5Anthropic | 55.3 | $4.00 |
| 12Claude Opus 4.7Anthropic | 55 | $10.00 |
| 13Muse Spark 1.1Meta | 53.2 | $2.00 |
| 14GPT-5.4OpenAI | 53.1 | $5.62 |
| 15GLM-5.2Z.AI +2 more hosts | 52.6 | from $2.15 |
| 16GPT-5.6 LunaOpenAI | 52.3 | $0.450 |
| 17Gemini 3.5 FlashGoogle | 52 | $3.38 |
| 18DeepSeek V4 FlashDeepSeek +2 more hosts | 51.8 | from $0.110 |
| 19Gemini 3.6 FlashGoogle | 51.6 | $3.00 |
| 20Claude Sonnet 4.6Anthropic | 48.4 | $6.00 |
LMSYS Arena ELO
Crowd head-to-head preference rating from blind human votes. Typically 1300–1550; higher is better.
| Model | Arena ELO | Cost $/M |
|---|---|---|
| 1GPT-5.5 ProOpenAI | 1510 | $67.50 |
| 2Claude Fable 5Anthropic | 1507 | $20.00 |
| 3Claude Opus 4.6Anthropic | 1498 | $10.00 |
| 4Qwen3.8-MaxAlibaba | 1497 | $3.00 |
| 5Claude Opus 4.7Anthropic | 1493 | $10.00 |
| 6Muse Spark 1.1Meta | 1488 | $2.00 |
| 7Gemini 3.1 ProGoogle | 1487 | $4.50 |
| 8Kimi K3Moonshot | 1485 | $6.00 |
| 9Gemini 3.6 FlashGoogle | 1485 | $3.00 |
| 10GPT-5.6 SolOpenAI | 1482 | $11.25 |
| 11GPT-5.4 ProOpenAI | 1478 | $67.50 |
| 12GPT-5.5OpenAI | 1477 | $11.25 |
| 13Gemini 3.5 FlashGoogle | 1476 | $3.38 |
| 14GPT-5.2 ChatOpenAI | 1476 | $4.81 |
| 15Grok 4.20xAI | 1475 | $1.56 |
| 16Gemini 3 Flash PreviewGoogle | 1473 | $1.12 |
| 17Claude Opus 4.8Anthropic | 1473 | $10.00 |
| 18Claude Sonnet 4.6Anthropic | 1472 | $6.00 |
| 19GLM-5.2Z.AI | 1469 | $2.15 |
| 20Grok 4.5xAI | 1468 | $3.00 |
These two indices use different scales and can't be added together — a model can rank highly on one and be absent from the other. Composite scores complement, and don't replace, the per-benchmark accuracy tables below.
Best reasoning — GPQA Diamond
PhD-level science reasoning. The hardest general reasoning benchmark.
| Model | GPQA score | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Sonnet 5Anthropic | 96.2% | $4.00 | 20.4 |
| 2GPT-5.6 SolOpenAI | 94.6% | $11.25 | 7.4 |
| 3GPT-5.4 ProOpenAI | 94.4% | $67.50 | — |
| 4Gemini 3.1 ProGoogle | 94.3% | $4.50 | 16.6 |
| 5Claude Opus 4.7Anthropic | 94.2% | $10.00 | 8.3 |
| 6Claude Mythos 5Anthropic | 94.1% | $20.00 | 4.4 |
| 7Claude Fable 5Anthropic | 94.1% | $20.00 | 4.5 |
| 8Claude Opus 4.8Anthropic | 93.6% | $10.00 | 8.1 |
| 9GPT-5.5OpenAI | 93.6% | $11.25 | 6.7 |
| 10Kimi K3Moonshot | 93.5% | $6.00 | 13.9 |
| 11MiniMax-M3MiniMax +2 more hosts | 93% | from $0.520 | 120.2 |
| 12GPT-5.6 TerraOpenAI | 92.9% | $4.50 | 19.5 |
| 13Qwen3.8-MaxAlibaba | 92.6% | $3.00 | — |
| 14GPT-5.2OpenAI | 92.4% | $4.81 | 15.8 |
| 15GPT-5.6 LunaOpenAI | 92.3% | $0.450 | 201.6 |
Best coding — SWE-Bench
Real-world software engineering tasks — resolving GitHub issues across popular repos.
| Model | SWE-Bench | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Opus 5Anthropic | 97% | $10.00 | 8.4 |
| 2GPT-5.6 SolOpenAI | 96.2% | $11.25 | 7.4 |
| 3Claude Mythos 5Anthropic | 95.5% | $20.00 | 4.4 |
| 4Claude Fable 5Anthropic | 95% | $20.00 | 4.5 |
| 5Kimi K3Moonshot | 93.4% | $6.00 | 13.9 |
| 6GPT-5.6 LunaOpenAI | 93% | $0.450 | 201.6 |
| 7Claude Opus 4.8Anthropic | 88.6% | $10.00 | 8.1 |
| 8Claude Opus 4.7Anthropic | 87.6% | $10.00 | 8.3 |
| 9Claude Sonnet 5Anthropic | 85.2% | $4.00 | 20.4 |
| 10GPT-5.3 CodexOpenAI | 85% | $4.81 | 14.8 |
| 11Claude Sonnet 4.5Anthropic | 82% | $6.00 | 12.5 |
| 12DeepSeek V4 ProDeepSeek +2 more hosts | 80.6% | from $0.540 | 135.2 |
| 13GPT-5.2OpenAI | 80% | $4.81 | 15.8 |
| 14DeepSeek V4 FlashDeepSeek +2 more hosts | 79% | from $0.110 | 442.9 |
| 15Gemini 3 Flash PreviewGoogle | 78% | $1.12 | — |
Terminal-Bench 2.1
Multi-step tasks in a terminal environment.
| Model | Terminal-Bench 2.1 | Cost $/M | Perf/$ |
|---|---|---|---|
| 1GPT-5.6 SolOpenAI | 88.8% | $11.25 | 7.4 |
| 2Kimi K3Moonshot | 88.3% | $6.00 | 13.9 |
| 3Claude Mythos 5Anthropic | 88% | $20.00 | 4.4 |
| 4Claude Fable 5Anthropic | 84.3% | $20.00 | 4.5 |
| 5GPT-5.6 LunaOpenAI | 84.3% | $0.450 | 201.6 |
| 6Grok 4.5xAI | 83.3% | $3.00 | — |
| 7GPT-5.5OpenAI | 82.7% | $11.25 | 6.7 |
| 8GPT-5.6 TerraOpenAI | 82.5% | $4.50 | 19.5 |
| 9GLM-5.2Z.AI +2 more hosts | 81% | from $2.15 | 35.3 |
| 10Claude Sonnet 5Anthropic | 80.4% | $4.00 | 20.4 |
| 11GPT-5.3 CodexOpenAI | 77.3% | $4.81 | 14.8 |
| 12Gemini 3.5 FlashGoogle | 76.2% | $3.38 | 20.8 |
| 13Claude Opus 4.8Anthropic | 74.6% | $10.00 | 8.1 |
| 14Gemini 3.1 ProGoogle | 70.3% | $4.50 | 16.6 |
| 15Claude Opus 4.7Anthropic | 69.4% | $10.00 | 8.3 |
LiveCodeBench
Contamination-resistant code generation.
| Model | LiveCodeBench | Cost $/M | Perf/$ |
|---|---|---|---|
| 1DeepSeek V4 ProDeepSeek +2 more hosts | 93.5% | from $0.540 | 135.2 |
| 2DeepSeek V4 FlashDeepSeek +2 more hosts | 91.6% | from $0.110 | 442.9 |
| 3Kimi K2.5Moonshot +1 more host | 85% | from $1.20 | 52.4 |
| 4Grok 4xAI | 79% | $6.00 | 10.5 |
| 5Claude Opus 4.6Anthropic | 76% | $10.00 | 7.7 |
| 6o3-miniOpenAI | 74.1% | $1.93 | 34.2 |
| 7Claude Sonnet 4.6Anthropic | 72.4% | $6.00 | 11.7 |
| 8Gemini 2.5 ProGoogle | 69% | $3.44 | 17.8 |
| 9GPT-OSS 20BFireworks +1 more host | 69% | from $0.130 | 490.2 |
| 10GPT-OSS 120BFireworks +2 more hosts | 69% | from $0.260 | 249.5 |
| 11Gemini 2.5 FlashGoogle | 63.5% | $0.850 | 59.1 |
| 12GPT-4.1OpenAI | 52% | $3.50 | 16.3 |
| 13Llama 4 MaverickDeepInfra | 41% | $0.350 | 186 |
| 14Llama 4 ScoutMeta +2 more hosts | 32.8% | from $0.150 | 279.4 |
MCP Atlas
Tool use over Model Context Protocol servers.
| Model | MCP Atlas | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Kimi K3Moonshot | 84.2% | $6.00 | 13.9 |
| 2Gemini 3.5 FlashGoogle | 83.6% | $3.38 | 20.8 |
| 3Claude Fable 5Anthropic | 83.3% | $20.00 | 4.5 |
| 4Claude Opus 4.8Anthropic | 82.2% | $10.00 | 8.1 |
| 5Claude Opus 4.7Anthropic | 79.1% | $10.00 | 8.3 |
| 6GLM-5.2Z.AI +2 more hosts | 77% | from $2.15 | 35.3 |
| 7Claude Opus 4.6Anthropic | 76.8% | $10.00 | 7.7 |
| 8GPT-5.5OpenAI | 75.3% | $11.25 | 6.7 |
| 9MiniMax-M3MiniMax +2 more hosts | 74.2% | from $0.520 | 120.2 |
| 10DeepSeek V4 ProDeepSeek +2 more hosts | 73.6% | from $0.540 | 135.2 |
| 11Claude Opus 4.5Anthropic | 69.8% | $10.00 | 7.1 |
| 12Claude Sonnet 4.6Anthropic | 69.5% | $6.00 | 11.7 |
| 13Gemini 3.1 ProGoogle | 69.2% | $4.50 | 16.6 |
| 14DeepSeek V4 FlashDeepSeek +2 more hosts | 69% | from $0.110 | 442.9 |
| 15GPT-5.2OpenAI | 67.6% | $4.81 | 15.8 |
BrowseComp
Agentic web search and information extraction.
| Model | BrowseComp | Cost $/M | Perf/$ |
|---|---|---|---|
| 1GPT-5.6 SolOpenAI | 92.2% | $11.25 | 7.4 |
| 2Kimi K3Moonshot | 91.2% | $6.00 | 13.9 |
| 3Claude Opus 5Anthropic | 90.8% | $10.00 | 8.4 |
| 4Claude Fable 5Anthropic | 88% | $20.00 | 4.5 |
| 5Gemini 3.1 ProGoogle | 85.9% | $4.50 | 16.6 |
| 6DeepSeek V4 FlashDeepSeek +2 more hosts | 85.9% | from $0.110 | 442.9 |
| 7Claude Sonnet 5Anthropic | 84.7% | $4.00 | 20.4 |
| 8GPT-5.5OpenAI | 84.4% | $11.25 | 6.7 |
| 9MiniMax-M3MiniMax +2 more hosts | 83.5% | from $0.520 | 120.2 |
| 10DeepSeek V4 ProDeepSeek +2 more hosts | 83.4% | from $0.540 | 135.2 |
| 11Kimi K2.6Moonshot +1 more host | 83.2% | from $1.71 | 44.5 |
| 12Claude Sonnet 4.6Anthropic | 76.2% | $6.00 | 11.7 |
OSWorld-Verified
Computer use — real GUI tasks on a desktop OS.
| Model | OSWorld-Verified | Cost $/M | Perf/$ |
|---|---|---|---|
| 1Claude Fable 5Anthropic | 85% | $20.00 | 4.5 |
| 2Claude Opus 4.8Anthropic | 83.4% | $10.00 | 8.1 |
| 3Claude Sonnet 5Anthropic | 81.2% | $4.00 | 20.4 |
| 4GPT-5.5OpenAI | 78.7% | $11.25 | 6.7 |
| 5Claude Sonnet 4.6Anthropic | 78.5% | $6.00 | 11.7 |
| 6Gemini 3.5 FlashGoogle | 78.4% | $3.38 | 20.8 |
| 7Claude Opus 4.7Anthropic | 78% | $10.00 | 8.3 |
| 8Gemini 3.1 ProGoogle | 76.2% | $4.50 | 16.6 |
| 9Kimi K2.6Moonshot +1 more host | 73.1% | from $1.71 | 44.5 |
| 10MiniMax-M3MiniMax +2 more hosts | 70.1% | from $0.520 | 120.2 |
| 11GPT-5.3 CodexOpenAI | 64.7% | $4.81 | 14.8 |
| 12GPT-5.6 SolOpenAI | 62.6% | $11.25 | 7.4 |
Best math — AIME 2025
American Invitational Mathematics Examination problems — competition-level math.
| Model | AIME score | Cost $/M | Perf/$ |
|---|---|---|---|
| 1GPT-5.2OpenAI | 100% | $4.81 | 15.8 |
| 2GPT-5.2 ProOpenAI | 100% | $57.75 | — |
| 3Claude Opus 4.6Anthropic | 99.8% | $10.00 | 7.7 |
| 4GPT-OSS 20BFireworks +1 more host | 98.7% | from $0.130 | 490.2 |
| 5GPT-OSS 120BFireworks +2 more hosts | 97.9% | from $0.260 | 249.5 |
| 6Claude Haiku 4.5Anthropic | 96.3% | $2.00 | 32.8 |
| 7Kimi K2.5Moonshot +1 more host | 96.1% | from $1.20 | 52.4 |
| 8GPT-5.4OpenAI | 95.5% | $5.62 | 14 |
| 9o4-miniOpenAI | 92.7% | $1.93 | 35.9 |
| 10Grok 4xAI | 91.7% | $6.00 | 10.5 |
| 11Gemini 2.5 ProGoogle | 88% | $3.44 | 17.8 |
| 12Claude Sonnet 4.5Anthropic | 87% | $6.00 | 12.5 |
| 13o3-miniOpenAI | 86.5% | $1.93 | 34.2 |
| 14Grok 4.3xAI | 85% | $1.56 | 46.3 |
| 15Claude Sonnet 4.6Anthropic | 83% | $6.00 | 11.7 |
Fastest endpoints — inference speed
Tokens per second on hosted inference — the one board here ranked by endpoint rather than by model, because speed is a property of the host's serving stack, not of the weights. The same model can appear more than once at different hosts, and the difference between those rows is the point. Speed matters for real-time applications and high-throughput pipelines.
| Model | Tokens/sec | Cost $/M |
|---|---|---|
| 1Llama 4 ScoutMeta | 2600 t/s | $0.170 |
| 2Llama 4 ScoutGroq | 2600 t/s | $0.170 |
| 3Llama 4 ScoutDeepInfra | 2600 t/s | $0.150 |
| 4Llama 3.3 70BMeta | 2500 t/s | $0.640 |
| 5Llama 3.3 70B VersatileGroq | 2500 t/s | $0.640 |
| 6Llama 3.3 70BTogether | 2500 t/s | $1.04 |
| 7GPT-OSS 120BFireworks | 1828.8 t/s | $0.260 |
| 8Llama 3.1 8BMeta | 1800 t/s | $0.060 |
| 9Llama 3.1 8B InstantGroq | 1800 t/s | $0.060 |
| 10GPT-OSS 20BFireworks | 941.5 t/s | $0.130 |
| 11Gemini 3.1 Flash-LiteGoogle | 600 t/s | $0.560 |
| 12GPT-OSS 20BGroq | 564 t/s | $0.130 |
| 13Qwen 3.6 27BGroq | 472.4 t/s | $1.20 |
| 14Gemini 2.5 Flash-LiteGoogle | 450 t/s | $0.180 |
| 15Granite 4 H SmallIBM | 414.5 t/s | $0.110 |
Hardest benchmark — Humanity's Last Exam
Expert-level questions across 100+ fields. No model scores above 65% — the frontier of AI capability.
| Model | HLE score | Cost $/M |
|---|---|---|
| 1Claude Opus 5Anthropic | 64.7% | $10.00 |
| 2Claude Mythos 5Anthropic | 64.5% | $20.00 |
| 3Claude Opus 4.8Anthropic | 57.9% | $10.00 |
| 4Claude Sonnet 5Anthropic | 57.4% | $4.00 |
| 5Kimi K3Moonshot | 56% | $6.00 |
| 6GLM-5.2Z.AI +2 more hosts | 54.7% | from $2.15 |
| 7Kimi K2.6Moonshot +1 more host | 54% | from $1.71 |
| 8DeepSeek V4 FlashDeepSeek +2 more hosts | 51.6% | from $0.110 |
| 9DeepSeek V4 ProDeepSeek +2 more hosts | 48.2% | from $0.540 |
| 10GPT-5.6 SolOpenAI | 47.2% | $11.25 |
Methodology
how the scores and rankings are computedBenchmark sources
Vellum LLM Leaderboard, Artificial Analysis, official model technical reports, and HuggingFace Open LLM Leaderboard. Scores are the latest publicly reported results as of August 2026, refreshed as new evaluations are published.
Performance per dollar
Shown two ways. The raw Perf/$ column is avg_benchmark_score / blended_cost_per_mtok (blended cost = (3×input + 1×output) ÷ 4 per million tokens). The Best Value ranking is driven instead by the comparability-normalized Norm Value: each model's benchmark strength is first converted to a percentile rank against the whole tracked population (avg_method = percentile_rank_vs_population), then divided by the same blended cost — so a model tested only on easy benchmarks can't post an inflated ratio.
A Norm Value is a percentile-derived value figure, not a % accuracy — only the raw Avg Score and per-benchmark columns are accuracy percentages.
Comparability rule
A model's "average" only covers the benchmarks it has actually been tested on, and the tracked benchmarks differ wildly in difficulty (nobody scores above ~65% on Humanity's Last Exam; AIME scores run into the high 90s), so averaging over different subsets isn't apples-to-apples. Every avg-based ranking and the scatter above therefore require at least 3 filled benchmarks and display the benchmark count (n). Single-benchmark tables (GPQA, SWE-Bench, Terminal-Bench, AIME, HLE and the rest) are unaffected — one shared benchmark is directly comparable.
One row per model
Open-weight models are often sold by several hosts at once — GLM-5.2 by Z.AI, Fireworks and Together. A benchmark score belongs to the weights, so all three score identically, and listing them separately would spend three ranking slots on one answer. Every board here shows a model once, at its cheapest host, with the remaining hosts noted on the row; the price reads “from” whenever more than one host sells it. The counts above follow the same rule — a model sold by three hosts counts once.
The one exception is Fastest endpoints. Tokens per second is a property of the host's serving stack rather than of the weights — Z.AI runs GLM-5.2 at a different speed from Fireworks — so that board ranks endpoints, and a model can legitimately appear on it more than once.
What each benchmark measures
- GPQA Diamond
- PhD-level questions in biology, chemistry, and physics
- AIME 2025
- American Invitational Mathematics Examination (competition math)
- SWE-Bench
- Resolving real GitHub issues in popular Python repositories
- Humanity's Last Exam
- Expert-level questions across 100+ academic fields
- ARC-AGI 2
- Visual reasoning and abstract pattern completion
- MMMLU
- Multilingual version of MMLU (57 subjects, multiple languages)
- HumanEval
- Python code generation from natural language descriptions
- MATH 500
- Mathematical problem solving across 5 difficulty levels
- BFCL
- Berkeley Function Calling Leaderboard (tool use & function calling)
Access this data programmatically via the benchmarks API endpoint or the API documentation.