Fastest LLM APIs for Low-Latency Apps
LLM APIs optimized for fast inference and low latency, ranked by measured throughput (tokens/sec). Compare verified pricing on fast inference providers like Groq, Together, and Fireworks.
MODELPRICEWATCH.COM · 2026-09-23
Cost calculator for this use casemonthly cost, top 3 models
- 1 Llama 4 Scout
- $—
- 2 Llama 3.3 70B
- $—
- 3 GPT-OSS 120B
- $—
Full ranking — top 26 models
list prices, USD per 1M tokens| Model | Throughput | Blended* | Input | Output | Context | Provider |
|---|---|---|---|---|---|---|
| 1Llama 4 Scout | 2600 tok/s | $0.168 | $0.110 | $0.340 | 10M | Meta |
| 2Llama 3.3 70B | 2500 tok/s | $0.640 | $0.590 | $0.790 | 128K | Meta |
| 3GPT-OSS 120B | 1829 tok/s | $0.262 | $0.150 | $0.600 | 128K | Fireworks |
| 4Llama 3.1 8B | 1800 tok/s | $0.058 | $0.050 | $0.080 | 128K | Meta |
| 5GPT-OSS 20B | 942 tok/s | $0.128 | $0.070 | $0.300 | 128K | Fireworks |
| 6Qwen3.6-27B | 472 tok/s | $1.20 | $0.600 | $3.00 | 131K | Groq |
| 7Gemini 3.5 Flash-Lite | 365 tok/s | $0.850 | $0.300 | $2.50 | 1M | |
| 8GLM-5.2 | 347 tok/s | $2.15 | $1.40 | $4.40 | 1M | Z.AI |
| 9Kimi K2.6 | 343 tok/s | $1.71 | $0.950 | $4.00 | 262K | Moonshot |
| 10Kimi K2.5 | 338 tok/s | $1.20 | $0.600 | $3.00 | 262K | Moonshot |
| 11Nova Micro | 300 tok/s | $0.061 | $0.035 | $0.140 | 128K | Amazon |
| 12Gemini 3.7 Flash | 288 tok/s | $1.50 | $0.750 | $3.75 | 1M | |
| 13Gemini 3.8 Flash | 278 tok/s | $1.50 | $0.750 | $3.75 | 1M | |
| 14Muse Spark 1.3 | 228 tok/s | $2.00 | $1.25 | $4.25 | 1M | Meta |
| 15Muse Spark 1.1 | 206 tok/s | $2.00 | $1.25 | $4.25 | 1M | Meta |
| 16Nova Lite | 200 tok/s | $0.105 | $0.060 | $0.240 | 300K | Amazon |
| 17DeepSeek V4 Flash | 200 tok/s | $0.175 | $0.140 | $0.280 | 1M | Fireworks |
| 18Codestral 2508 | 200 tok/s | $0.450 | $0.300 | $0.900 | 256K | Mistral |
| 19GPT-4.1 mini | 200 tok/s | $0.700 | $0.400 | $1.60 | 1M | OpenAI |
| 20GPT-5.4 nano | 191 tok/s | $0.463 | $0.200 | $1.25 | 400K | OpenAI |
| 21GLM-5.3 | 72 tok/s | $2.15 | $1.40 | $4.40 | 1M | Z.AI |
| 22Claude Fable 5.1 | 65 tok/s | $20.00 | $10.00 | $50.00 | 1M | Anthropic |
| 23Qwen3.8-27B | 44 tok/s | $1.13 | $0.500 | $3.00 | 1M | Alibaba |
| 24Qwen3.8-Flash | — | $0.230 | $0.150 | $0.470 | 1M | Alibaba |
| 25Mercury 2.5 | — | $0.338 | $0.200 | $0.750 | 260K | Inception |
| 26GLM-5.3-FlashX | — | $0.590 | $0.370 | $1.25 | 1M | Z.AI |
How models are selected
Generally-available models with independently measured throughput (output tokens per second) or a speed-optimized release, ranked by measured throughput. One row per model — host duplicates are collapsed. Models without a speed measurement rank below measured ones, cheapest first.
Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.