NVIDIA Nemotron 3.5 Lightning vs DeepSeek V4 Flash
Side-by-side comparison of API pricing, specs, benchmarks, and capabilities
| Specification |
by DeepInfra
|
by DeepInfra
3 providers:
DeepSeek $0.440/$1.32
Fireworks $0.140/$0.280
DeepInfra $0.090/$0.180
|
|---|---|---|
| Overview | ||
| Status | Current open weights Open weights | Current budget |
| Released | Aug 11, 2026 | Apr 24, 2026 |
| Pricing per million tokens | ||
| Input | $0.080/Mtok | $0.090/Mtok |
| Output | $0.200/Mtok | $0.180/Mtok |
| Blended avg | $0.110/Mtok | $0.113/Mtok |
| Cached input | $0.040/Mtok | $0.018/Mtok |
| Price basis | — | DeepInfra lists two entries for this model and prices them differently: the floating alias DeepSeek-V4-Flash at $0.09 input / $0.018 cached input / $0.18 output per 1M tokens, and the pinned snapshot DeepSeek-V4-Flash-0731 at $0.08 / $0.016 / $0.18. The tracked figure is the floating alias, because that is what a caller who names the model without a date gets. Pinning to the 0731 snapshot saves 11% on input at the cost of being left behind when DeepInfra moves the alias. |
| Specifications | ||
| Context window | 256K tokens | 1M tokens |
| Parameters | 30B (3B active) | Proprietary |
| Speed (TPS) | — | — |
| Modalities | ||
| Input |
text
|
text
|
| Benchmarks sources: Artificial Analysis / DeepSeek model card | ||
| Independent composite scores measured for both models — each on its own scale | ||
| AA Intelligence Index | 23.6 | 51.8 |
| AA Agentic Index | 13.8 | 48.4 |
| AA-Omniscience Index (−100–100) | -17.7 | -14.3 |
| GDPval-AA v2 | — | 1558.4 ELO |
| Per-benchmark results published for DeepSeek V4 Flash only — no independent per-benchmark scores exist for NVIDIA Nemotron 3.5 Lightning yet | ||
| Avg benchmark score | — | 77.5 |
| Perf / dollar | — | 688.9 |
| GPQA Diamond | — | 88.1 |
| SWE-Bench Verified | — | 79 vendor |
| LiveCodeBench | — | 91.6 |
| MCP Atlas | — | 69 |
| BrowseComp | — | 85.9 |
| Humanity's Last Exam | — | 51.6 |
| ARC-AGI 2 | — | 18 vendor |
| AIME 2025 | — | 58 vendor |
| MMMLU | — | 75 vendor |
| BFCL | — | 62 vendor |
| HumanEval | — | 79 |
| MATH 500 | — | 68 vendor |
| Providers | ||
| Available from |
DeepInfra — $0.080/$0.200/Mtok
|
DeepSeek — $0.440/$1.32/Mtok
Fireworks — $0.140/$0.280/Mtok
DeepInfra — $0.090/$0.180/Mtok
|
Cost at scale
1M tokens · 50/50 input/output| Volume | NVIDIA Nemotron 3.5 Lightning | DeepSeek V4 Flash | Savings |
|---|---|---|---|
| 1M tokens | $0.11 | $0.11 | $0 (0%) |
| 10M tokens | $1.1 | $1.13 | $0.02 (1.8%) |
| 100M tokens | $11 | $11.25 | $0.25 (2.2%) |
| 1000M tokens | $110 | $112.5 | $2.5 (2.2%) |
When to pick which
distilled from the pricing and spec data above- cost dominates: $0.110/Mtok blended vs $0.113 — 2.2% less on the same 50/50 token mix
- you want open weights — self-host it, fine-tune it, or exit the API entirely; DeepSeek V4 Flash is closed
- your workload re-reads context (agents, RAG, long chats): cached input costs $0.018/Mtok — 80% off its list input price
- you need the longer context: 1M tokens vs 256K (3.9×)
- provider flexibility: available from 3 providers, so you can price-shop hosts or fail over
Summary
NVIDIA Nemotron 3.5 Lightning by DeepInfra costs $0.080/Mtok input and $0.200/Mtok output, with a 256K-token context window. It supports text input.
DeepSeek V4 Flash by DeepInfra costs $0.090/Mtok input and $0.180/Mtok output, with a 1M-token context window. It supports text input and is available from 3 providers.
On a blended cost basis, NVIDIA Nemotron 3.5 Lightning is 2.2% cheaper than DeepSeek V4 Flash.
The two aren't directly comparable on average benchmark score: DeepSeek V4 Flash has published per-benchmark results, while NVIDIA Nemotron 3.5 Lightning does not yet — it is measured today on independent composites (see the table above).
Note: Pricing is per million tokens. Actual costs vary with usage patterns, prompt caching, and batch discounts. Always verify against official provider pricing pages.