ModelPriceWatch.com
Last scan 2026-08-09 Models tracked 198 Providers 30 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

Best LLM APIs for Reasoning Tasks

LLM APIs specialized for reasoning and complex problem solving. Models ranked by GPQA Diamond score, with verified pricing for math, logic, and multi-step inference.

68 models qualify top 25 shown ranked by GPQA Diamond
1Anthropic

Claude Sonnet 5

96.2% GPQA Diamond

$4.00/1M blended · $2.00 in · $10.00 out

2OpenAI

GPT-5.6 Sol

94.6% GPQA Diamond

$11.25/1M blended · $5.00 in · $30.00 out

3Google

Gemini 3.1 Pro

94.3% GPQA Diamond

$4.50/1M blended · $2.00 in · $12.00 out

MODELPRICEWATCH.COM · 2026-08-09

Cost calculator for this use casemonthly cost, top 3 models

1 Claude Sonnet 5
$—
2 GPT-5.6 Sol
$—
3 Gemini 3.1 Pro
$—

Full ranking — top 25 models

list prices, USD per 1M tokens
Top 25 models for Best LLM APIs for Reasoning Tasks, ranked by GPQA Diamond
Model GPQA Diamond Blended* Input Output Context Provider
1Claude Sonnet 5 96.2% $4.00 $2.00 $10.00 1M Anthropic
2GPT-5.6 Sol 94.6% $11.25 $5.00 $30.00 1M OpenAI
3Gemini 3.1 Pro 94.3% $4.50 $2.00 $12.00 2M Google
4Claude Opus 4.7 94.2% $10.00 $5.00 $25.00 1M Anthropic
5Claude Fable 5 94.1% $20.00 $10.00 $50.00 1M Anthropic
6Claude Opus 4.8 93.6% $10.00 $5.00 $25.00 1M Anthropic
7GPT-5.5 93.6% $11.25 $5.00 $30.00 1M OpenAI
8Kimi K3 93.5% $6.00 $3.00 $15.00 1M Moonshot
9MiniMax-M3 93% $0.525 $0.300 $1.20 1M MiniMax
10GPT-5.6 Terra 92.9% $4.50 $2.00 $12.00 1M OpenAI
11GPT-5.2 92.4% $4.81 $1.75 $14.00 400K OpenAI
12GPT-5.6 Luna 92.3% $0.450 $0.200 $1.20 1M OpenAI
13Claude Opus 4.6 91.3% $10.00 $5.00 $25.00 1M Anthropic
14GLM-5.2 91.2% $2.15 $1.40 $4.40 1M Z.AI
15Kimi K2.6 90.5% $1.71 $0.950 $4.00 262K Moonshot
16DeepSeek V4 Pro 90.1% $0.544 $0.435 $0.870 1M DeepSeek
17Claude Sonnet 4.6 89.9% $6.00 $3.00 $15.00 1M Anthropic
18DeepSeek V4 Flash 88.1% $0.175 $0.140 $0.280 1M DeepSeek
19GPT-5.4 88% $5.63 $2.50 $15.00 1M OpenAI
20Kimi K2.5 87.6% $1.20 $0.600 $3.00 262K Moonshot
21Qwen3.8-Max 92.6% vendor $3.00 $2.00 $6.00 1M Alibaba
22Inkling Small $0.525 $0.300 $1.20 262K Thinking Machines
23Inkling $1.76 $1.00 $4.05 262K Thinking Machines
24Muse Spark 1.2 $2.00 $1.25 $4.25 1M Meta
25Claude Opus 5 $10.00 $5.00 $25.00 1M Anthropic
Ranked by GPQA Diamond — independent leaderboard data; benchmark citations are on each model page. Scores flagged “vendor” are the maker's own claim, shown for reference and ranked below independently-scored models; models with no published score rank last, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.

Recent price movement in this ranking

price-only deltas · logged by the daily scan

3 of the top 25 reasoning models have re-priced since we began tracking · last checked Aug 9, 2026.

Models in this Best LLM APIs for Reasoning Tasks ranking that have re-priced since tracking began, most recent move first
Model Changes Latest move Provider
GPT-5.6 Terra 1 price cut on Aug 1, 2026: $2.5/$15 → $2/$12 /Mtok OpenAI
GPT-5.6 Luna 1 price cut on Aug 1, 2026: $1/$6 → $0.2/$1.2 /Mtok OpenAI
MiniMax-M3 1 price cut on Jun 17, 2026: $0.6/$2.4 → $0.3/$1.2 /Mtok MiniMax
Changes = distinct price moves logged since tracking began; in/out prices are $ per 1M tokens. MODELPRICEWATCH.COM · 2026-08-09

Most recently, GPT-5.6 Terra cut its price on Aug 1, 2026 — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →

How models are selected

Generally-available models with a GPQA Diamond score (graduate-level science questions) or a reasoning-focused release, ranked by independently-evaluated GPQA Diamond. One row per model — host duplicates are collapsed. Vendor-reported scores are shown flagged and ranked below independently-scored models, never against them; models with no published score rank last, cheapest first.

Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.

Other use case rankings