ModelPriceWatch.com
Last scan 2026-08-09 Models tracked 198 Providers 30 Cheapest paid Granite 4.0 Micro $0.017/Mtok in Every price links to its source

Fastest LLM APIs for Low-Latency Apps

LLM APIs optimized for fast inference and low latency, ranked by measured throughput (tokens/sec). Compare verified pricing on fast inference providers like Groq, Together, and Fireworks.

74 models qualify top 27 shown ranked by Throughput
1Meta

Llama 4 Scout

2600 tok/s Throughput

$0.168/1M blended · $0.110 in · $0.340 out

2Meta

Llama 3.3 70B

2500 tok/s Throughput

$0.640/1M blended · $0.590 in · $0.790 out

3Fireworks

GPT OSS 120B

1829 tok/s Throughput

$0.262/1M blended · $0.150 in · $0.600 out

MODELPRICEWATCH.COM · 2026-08-09

Cost calculator for this use casemonthly cost, top 3 models

1 Llama 4 Scout
$—
2 Llama 3.3 70B
$—
3 GPT OSS 120B
$—

Full ranking — top 27 models

list prices, USD per 1M tokens
Top 27 models for Fastest LLM APIs for Low-Latency Apps, ranked by Throughput
Model Throughput Blended* Input Output Context Provider
1Llama 4 Scout 2600 tok/s $0.168 $0.110 $0.340 10M Meta
2Llama 3.3 70B 2500 tok/s $0.640 $0.590 $0.790 128K Meta
3GPT OSS 120B 1829 tok/s $0.262 $0.150 $0.600 128K Fireworks
4Llama 3.1 8B 1800 tok/s $0.058 $0.050 $0.080 128K Meta
5GPT OSS 20B 942 tok/s $0.128 $0.070 $0.300 128K Fireworks
6Qwen 3.6 27B 472 tok/s $1.20 $0.600 $3.00 128K Groq
7Granite 4 H Small 415 tok/s $0.107 $0.060 $0.250 128K IBM
8GLM-5.2 347 tok/s $2.15 $1.40 $4.40 1M Z.AI
9Kimi K2.6 343 tok/s $1.71 $0.950 $4.00 262K Moonshot
10Kimi K2.5 338 tok/s $1.20 $0.600 $3.00 262K Moonshot
11Nova Micro 300 tok/s $0.061 $0.035 $0.140 128K Amazon
12Nova Lite 200 tok/s $0.105 $0.060 $0.240 300K Amazon
13DeepSeek V4 Flash 200 tok/s $0.175 $0.140 $0.280 1M Fireworks
14Granite 4 H Medium 200 tok/s $0.262 $0.150 $0.600 128K IBM
15Codestral 2508 200 tok/s $0.450 $0.300 $0.900 256K Mistral
16GPT-4.1 mini 200 tok/s $0.700 $0.400 $1.60 1M OpenAI
17GPT-5.4 nano 191 tok/s $0.463 $0.200 $1.25 400K OpenAI
18GPT-5.6 Luna 185 tok/s $0.450 $0.200 $1.20 1M OpenAI
19GPT-5.4 mini 180 tok/s $1.69 $0.750 $4.50 1M OpenAI
20Gemini 3.5 Flash 175 tok/s $3.38 $1.50 $9.00 1M Google
21GPT-5.6 Terra 136 tok/s $4.50 $2.00 $12.00 1M OpenAI
22Grok 4.5 88 tok/s $3.00 $2.00 $6.00 500K xAI
23GPT-5.6 Sol 54 tok/s $11.25 $5.00 $30.00 1M OpenAI
24Kimi K3 35 tok/s $6.00 $3.00 $15.00 1M Moonshot
25Qwen3.7-Flash $0.055 $0.030 $0.130 1M Alibaba
26Gemini 3.5 Flash-Lite $0.850 $0.300 $2.50 1M Google
27Gemini 3.6 Flash $3.00 $1.50 $7.50 1M Google
Ranked by Throughput — independent leaderboard data; benchmark citations are on each model page. Scores flagged “vendor” are the maker's own claim, shown for reference and ranked below independently-scored models; models with no published score rank last, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.

Recent price movement in this ranking

price-only deltas · logged by the daily scan

2 of the top 27 fast inference models have re-priced since we began tracking · last checked Aug 9, 2026.

Models in this Fastest LLM APIs for Low-Latency Apps ranking that have re-priced since tracking began, most recent move first
Model Changes Latest move Provider
GPT-5.6 Luna 1 price cut on Aug 1, 2026: $1/$6 → $0.2/$1.2 /Mtok OpenAI
GPT-5.6 Terra 1 price cut on Aug 1, 2026: $2.5/$15 → $2/$12 /Mtok OpenAI
Changes = distinct price moves logged since tracking began; in/out prices are $ per 1M tokens. MODELPRICEWATCH.COM · 2026-08-09

Most recently, GPT-5.6 Luna cut its price on Aug 1, 2026 — a sign pricing in this category is still moving, so re-check before committing to a long-term choice. See all recent price moves →

How models are selected

Generally-available models with independently measured throughput (output tokens per second) or a speed-optimized release, ranked by measured throughput. One row per model — host duplicates are collapsed. Models without a speed measurement rank below measured ones, cheapest first.

Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.

Other use case rankings