LIVE Cheapest paid: Granite 4.0 Micro $0.017/Mtok in 174 models tracked Updated Jul 26, 2026
Jul 26, 2026
ModelPriceWatch$/Mtok
Pricing / Compare / Llama 4 Scout vs Llama Nemotron Ultra 253B

Llama 4 Scout vs Llama Nemotron Ultra 253B

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

Llama 4 Scout is 90.5% cheaper on blended cost ($0.200 vs $2.10/Mtok)
 
3 providers: Groq $0.110/$0.340 Meta $0.110/$0.340 DeepInfra $0.100/$0.300
Nby NVIDIA
Overview
StatusCurrent open weights Open weights Current open weights Open weights
Released Apr 6, 2025 Jan 1, 2025
Pricing per million tokens
Input $0.100/Mtok $0.600/Mtok
Output $0.300/Mtok $3.60/Mtok
Blended avg $0.200/Mtok $2.10/Mtok
Specifications
Context window 10M tokens 128K tokens
Parameters 17B (16 experts) 253B
Speed (TPS)
Modalities
Input
textimage
text
Benchmarks sources: Vellum LLM Leaderboard (Jun 2026), Meta model card / Vellum LLM Leaderboard
Avg benchmark score 46.1 70
Perf / dollar 230.5 33.3
GPQA Diamond 50 76
SWE-Bench Verified 28 30 vendor
LiveCodeBench 32.8 64
Humanity's Last Exam 12 15 vendor
ARC-AGI 2 15 14 vendor
AIME 2025 45 45 vendor
MMMLU 75 75 vendor
BFCL 58 58 vendor
HumanEval 80 80 vendor
MATH 500 65 65 vendor
Independent composite scores each on its own scale — not part of the average above
LMSYS Chatbot Arena 1347.5 ELO
Providers
Available from
Groq — $0.110/$0.340/Mtok
Meta — $0.110/$0.340/Mtok
DeepInfra — $0.100/$0.300/Mtok
NVIDIA — $0.600/$3.60/Mtok

Cost at scale — 1M tokens (50/50 input/output)

VolumeLlama 4 ScoutLlama Nemotron Ultra 253BSavings
1M tokens $0.2 $2.1 $1.9 (90.5%)
10M tokens $2 $21 $19 (90.5%)
100M tokens $20 $210 $190 (90.5%)
1000M tokens $200 $2100 $1900 (90.5%)

Summary

Llama 4 Scout by DeepInfra costs $0.100/Mtok input and $0.300/Mtok output, with a 10M-token context window. It supports text, image input and is available from 3 providers.

Llama Nemotron Ultra 253B by NVIDIA costs $0.600/Mtok input and $3.60/Mtok output, with a 128K-token context window. It supports text input.

On a blended cost basis, Llama 4 Scout is 90.5% cheaper than Llama Nemotron Ultra 253B. It also has a larger context window.

On benchmarks, Llama Nemotron Ultra 253B scores higher (70 vs 46.1) on average. In terms of value, Llama 4 Scout has better performance per dollar (230.5 vs 33.3).

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.