Best LLM APIs for Reasoning Tasks
LLM APIs specialized for reasoning and complex problem solving. Models ranked by GPQA Diamond score, with verified pricing for math, logic, and multi-step inference.
MODELPRICEWATCH.COM · 2026-09-22
Cost calculator for this use casemonthly cost, top 3 models
- 1 Claude Sonnet 5
- $—
- 2 GPT-5.6 Sol
- $—
- 3 Gemini 3.1 Pro
- $—
Full ranking — top 27 models
list prices, USD per 1M tokens| Model | GPQA Diamond | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|---|
| 1Claude Sonnet 5 | 96.2% | $4.00 | $2.00 | $10.00 | 1M | Anthropic |
| 2GPT-5.6 Sol | 94.6% | $8.00 | $4.00 | $20.00 | 1M | OpenAI |
| 3Gemini 3.1 Pro | 94.3% | $4.50 | $2.00 | $12.00 | 2M | |
| 4Claude Opus 4.7 | 94.2% | $10.00 | $5.00 | $25.00 | 1M | Anthropic |
| 5Claude Fable 5 | 94.1% | $20.00 | $10.00 | $50.00 | 1M | Anthropic |
| 6Claude Opus 4.8 | 93.6% | $10.00 | $5.00 | $25.00 | 1M | Anthropic |
| 7GPT-5.5 | 93.6% | $11.25 | $5.00 | $30.00 | 1M | OpenAI |
| 8Kimi K3 | 93.5% | $6.00 | $3.00 | $15.00 | 1M | Moonshot |
| 9Claude Fable 5.1 | 93.4% | $20.00 | $10.00 | $50.00 | 1M | Anthropic |
| 10MiniMax-M3 | 93% | $0.525 | $0.300 | $1.20 | 1M | MiniMax |
| 11GPT-5.6 Terra | 92.9% | $4.50 | $2.00 | $12.00 | 1M | OpenAI |
| 12GPT-5.2 | 92.4% | $4.81 | $1.75 | $14.00 | 400K | OpenAI |
| 13GPT-5.6 Luna | 92.3% | $0.450 | $0.200 | $1.20 | 1M | OpenAI |
| 14Claude Opus 4.6 | 91.3% | $10.00 | $5.00 | $25.00 | 1M | Anthropic |
| 15GLM-5.2 | 91.2% | $2.15 | $1.40 | $4.40 | 1M | Z.AI |
| 16Kimi K2.6 | 90.5% | $1.71 | $0.950 | $4.00 | 262K | Moonshot |
| 17DeepSeek V4 Pro | 90.1% | $1.63 | $1.30 | $2.60 | 1M | DeepInfra |
| 18Claude Sonnet 4.6 | 89.9% | $6.00 | $3.00 | $15.00 | 1M | Anthropic |
| 19DeepSeek V4 Flash | 88.1% | $0.113 | $0.090 | $0.180 | 1M | DeepInfra |
| 20GPT-5.4 | 88% | $5.63 | $2.50 | $15.00 | 1M | OpenAI |
| 21Qwen3.8-27B | 89.2% vendor | $1.13 | $0.500 | $3.00 | 1M | Alibaba |
| 22GLM-5.3-Flash | — | $0.237 | $0.150 | $0.500 | 1M | Z.AI |
| 23Mercury 2.5 | — | $0.338 | $0.200 | $0.750 | 260K | Inception |
| 24Muse Spark 1.3 | — | $2.00 | $1.25 | $4.25 | 1M | Meta |
| 25GLM-5.3 | — | $2.15 | $1.40 | $4.40 | 1M | Z.AI |
| 26Fugu Max | — | $3.00 | $2.00 | $6.00 | 1M | Sakana AI |
| 27Claude Opus 5.5 | — | $8.00 | $4.00 | $20.00 | 1M | Anthropic |
Recent price movement in this ranking
price-only deltas · logged by the daily scan5 of the top 27 reasoning models have re-priced since we began tracking · last checked .
| Model | Changes | Latest move | Provider |
|---|---|---|---|
| GLM-5.3-Flash | 1 | price rise on : $0.075/$0.250 → $0.150/$0.500 /Mtok | Z.AI |
| GPT-5.6 Sol | 1 | price cut on : $5.00/$30.00 → $4.00/$20.00 /Mtok | OpenAI |
| GPT-5.6 Terra | 1 | price cut on : $2.50/$15.00 → $2.00/$12.00 /Mtok | OpenAI |
| GPT-5.6 Luna | 1 | price cut on : $1.00/$6.00 → $0.200/$1.20 /Mtok | OpenAI |
| MiniMax-M3 | 1 | price cut on : $0.600/$2.40 → $0.300/$1.20 /Mtok | MiniMax |
Most recently, GLM-5.3-Flash raised its price on — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →
How models are selected
Generally-available models with a GPQA Diamond score (graduate-level science questions) or a reasoning-focused release, ranked by independently-evaluated GPQA Diamond. One row per model — host duplicates are collapsed. Vendor-reported scores are shown flagged and ranked below independently-scored models, never against them; models with no published score rank last, cheapest first.
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.