ModelPriceWatch.com
Last scan 2026-08-28 Models tracked 237 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Llama 3.1 8B vs Llama 4 Maverick

Side-by-side comparison of API pricing, specs, benchmarks, and capabilities

Llama 3.1 8B is 83.6% cheaper on blended cost ($0.058 vs $0.350/Mtok)
Specification
by Meta
Overview
StatusCurrent open weights Open weights Current open weights Open weights
Released Jul 23, 2024 Apr 6, 2025
Pricing per million tokens
Input $0.050/Mtok $0.200/Mtok
Output $0.080/Mtok $0.800/Mtok
Blended avg $0.058/Mtok $0.350/Mtok
Specifications
Context window 128K tokens 1M tokens
Parameters 8B 17B (128 experts)
Speed (TPS)
Modalities
Input
text
textimage
Benchmarks sources: Vellum LLM Leaderboard (Jun 2026), Kaggle dataset / Vellum LLM Leaderboard
Avg benchmark score 34 65.1
Perf / dollar 591.3 186
GPQA Diamond 30 69.8
MMMLU 65 84.6
SWE-Bench Verified 12
LiveCodeBench 41
Humanity's Last Exam 5
ARC-AGI 2 6
AIME 2025 20
BFCL 48
HumanEval 72
MATH 500 48
Independent composite scores each on its own scale — not part of the average above
AA Intelligence Index 14.5
AA Agentic Index 1.2
AA-Omniscience Index (−100–100) -41.9
Providers
Available from
Meta — $0.050/$0.080/Mtok
DeepInfra — $0.200/$0.800/Mtok

Cost at scale

1M tokens · 50/50 input/output
Projected cost of Llama 3.1 8B vs Llama 4 Maverick at increasing token volumes
VolumeLlama 3.1 8BLlama 4 MaverickSavings
1M tokens $0.06 $0.35 $0.29 (82.9%)
10M tokens $0.58 $3.5 $2.93 (83.7%)
100M tokens $5.75 $35 $29.25 (83.6%)
1000M tokens $57.5 $350 $292.5 (83.6%)

When to pick which

distilled from the pricing and spec data above
Pick Llama 3.1 8B if…
  • cost dominates: $0.058/Mtok blended vs $0.350 — 83.6% less on the same 50/50 token mix
Pick Llama 4 Maverick if…
  • you need the longer context: 1M tokens vs 128K (7.8×)
  • raw capability matters more than price: 65.1 vs 34 average benchmark score

Summary

Llama 3.1 8B by Meta costs $0.050/Mtok input and $0.080/Mtok output, with a 128K-token context window. It supports text input.

Llama 4 Maverick by DeepInfra costs $0.200/Mtok input and $0.800/Mtok output, with a 1M-token context window. It supports text, image input.

On a blended cost basis, Llama 3.1 8B is 83.6% cheaper than Llama 4 Maverick.

On benchmarks, Llama 4 Maverick scores higher (65.1 vs 34) on average. In terms of value, Llama 3.1 8B has better performance per dollar (591.3 vs 186).

Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.

More comparisons

pairs sharing a model with this page