ModelPriceWatch.com
Last scan 2026-09-15 Models tracked 259 Providers 32 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

Llama 3.1 8B Instant

by Groq · 8B parameters

Retired fast open weights 840 TPS cheap tier
Retired on . No longer served on the endpoint we price — the figures below are kept for historical reference only. Replaced by GPT-OSS 20B.
Availability: Groq announced this model's shutdown on 2026-06-17 and shut it down on 2026-08-16, naming GPT-OSS 20B as the replacement. It is gone from Groq's models page. Groq scopes the deprecation — "this deprecation applies to free and developer-tier usage; enterprise customers with a committed-spend contract are not affected" — so the model may still be served under a negotiated enterprise contract. The rate above is the public per-token price it last carried, which is the price this site tracks and one you can no longer buy at.

Today's price · per 1M tokens

Input

$0.050

per 1M tokens

Output

$0.080

per 1M tokens

Blended

$0.058

blended $/1M — 3:1 weighted input:output

Source: official Groq pricing · read MODELPRICEWATCH.COM · 2026-09-15

Price receipt

We read Groq's own pricing page on and found Llama 3.1 8B Instant 128k listed with a price on the same row — that read is where the number above comes from, recorded as a correction. We keep the page text we read; its content hash is 2edf4f7dbe.

Overview

Retired from Groq on 2026-08-16; the endpoint we price no longer serves it. Llama 3.1 8B on Groq's LPU, last listed at $0.05/$0.08 per 1M tokens and about 840 tokens/sec. Groq names GPT-OSS 20B as the replacement.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 1/5
Text Image Audio Video PDF
Output 1/5
Text Image Audio Video Embedding
Features 3/9
Prompt caching Reasoning Coding Fast inference Long context Open weights Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Avg benchmark score34%
Perf per $/Mtok591.3
GPQA Diamond
30%
SWE-Bench Verified
12%
Humanity's Last Exam
5%
ARC-AGI 2
6%
AIME 2025
20%
MMMLU
65%
BFCL
48%
HumanEval
72%
MATH 500
48%
How it stacks up
  • Ranks #83 of 84 benchmarked models by average score
  • Ranks #170 of 181 comparably-measured models by percentile score, across 9 independent measurements
  • Ranks #40 of 181 comparably-tested models by normalized performance per dollar
  • Strongest at HumanEval — 72%, #43 of 49

1800 tokens/sec output

Source: Vellum LLM Leaderboard (Jun 2026), Kaggle dataset · updated · See full rankings →

Specifications

Provider
Groq
Context window
128K tokens
Modality
text
Parameters
8B
Open source
Yes — open weights available
Released
Status
Retired ended
Last updated
Tags
fastopen-weightsspeedretired

Availability verified: per Groq's own deprecation notice