ModelPriceWatch.com
Last scan 2026-09-23 Models tracked 267 Providers 35 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Fastest LLM APIs for Low-Latency Apps

LLM APIs optimized for fast inference and low latency, ranked by measured throughput (tokens/sec). Compare verified pricing on fast inference providers like Groq, Together, and Fireworks.

90 models qualify top 26 shown ranked by Throughput
1Meta

Llama 4 Scout

2600 tok/s Throughput

$0.168/1M blended · $0.110 in · $0.340 out

2Meta

Llama 3.3 70B

2500 tok/s Throughput

$0.640/1M blended · $0.590 in · $0.790 out

3Fireworks

GPT-OSS 120B

1829 tok/s Throughput

$0.262/1M blended · $0.150 in · $0.600 out

MODELPRICEWATCH.COM · 2026-09-23

Cost calculator for this use casemonthly cost, top 3 models

1 Llama 4 Scout
$—
2 Llama 3.3 70B
$—
3 GPT-OSS 120B
$—

Full ranking — top 26 models

list prices, USD per 1M tokens
Top 26 models for Fastest LLM APIs for Low-Latency Apps, ranked by Throughput
Model Throughput Blended* Input Output Context Provider
1Llama 4 Scout 2600 tok/s $0.168 $0.110 $0.340 10M Meta
2Llama 3.3 70B 2500 tok/s $0.640 $0.590 $0.790 128K Meta
3GPT-OSS 120B 1829 tok/s $0.262 $0.150 $0.600 128K Fireworks
4Llama 3.1 8B 1800 tok/s $0.058 $0.050 $0.080 128K Meta
5GPT-OSS 20B 942 tok/s $0.128 $0.070 $0.300 128K Fireworks
6Qwen3.6-27B 472 tok/s $1.20 $0.600 $3.00 131K Groq
7Gemini 3.5 Flash-Lite 365 tok/s $0.850 $0.300 $2.50 1M Google
8GLM-5.2 347 tok/s $2.15 $1.40 $4.40 1M Z.AI
9Kimi K2.6 343 tok/s $1.71 $0.950 $4.00 262K Moonshot
10Kimi K2.5 338 tok/s $1.20 $0.600 $3.00 262K Moonshot
11Nova Micro 300 tok/s $0.061 $0.035 $0.140 128K Amazon
12Gemini 3.7 Flash 288 tok/s $1.50 $0.750 $3.75 1M Google
13Gemini 3.8 Flash 278 tok/s $1.50 $0.750 $3.75 1M Google
14Muse Spark 1.3 228 tok/s $2.00 $1.25 $4.25 1M Meta
15Muse Spark 1.1 206 tok/s $2.00 $1.25 $4.25 1M Meta
16Nova Lite 200 tok/s $0.105 $0.060 $0.240 300K Amazon
17DeepSeek V4 Flash 200 tok/s $0.175 $0.140 $0.280 1M Fireworks
18Codestral 2508 200 tok/s $0.450 $0.300 $0.900 256K Mistral
19GPT-4.1 mini 200 tok/s $0.700 $0.400 $1.60 1M OpenAI
20GPT-5.4 nano 191 tok/s $0.463 $0.200 $1.25 400K OpenAI
21GLM-5.3 72 tok/s $2.15 $1.40 $4.40 1M Z.AI
22Claude Fable 5.1 65 tok/s $20.00 $10.00 $50.00 1M Anthropic
23Qwen3.8-27B 44 tok/s $1.13 $0.500 $3.00 1M Alibaba
24Qwen3.8-Flash $0.230 $0.150 $0.470 1M Alibaba
25Mercury 2.5 $0.338 $0.200 $0.750 260K Inception
26GLM-5.3-FlashX $0.590 $0.370 $1.25 1M Z.AI
Ranked by Throughput — independent leaderboard data; benchmark citations are on each model page. Scores flagged “vendor” are the maker's own claim, shown for reference and ranked below independently-scored models; models with no published score rank last, cheapest first. * Blended = (3×input + 1×output) ÷ 4. Every price links to its source on the model page.

How models are selected

Generally-available models with independently measured throughput (output tokens per second) or a speed-optimized release, ranked by measured throughput. One row per model — host duplicates are collapsed. Models without a speed measurement rank below measured ones, cheapest first.

Prices are per million tokens (Mtok), sourced directly from — and linked to — official provider pricing pages. "Blended cost" is (3×input + 1×output) ÷ 4 — weighted toward input because real workloads read far more tokens than they generate.

Other use case rankings