ModelPriceWatch.com
Last scan 2026-08-18 Models tracked 217 Providers 31 Cheapest paid Granite 4.0 H Micro $0.017/Mtok in Every price links to its source

NVIDIA Nemotron 3.5 Lightning

by DeepInfra · 30B (3B active) parameters

Current new · 7d open weights open weights cheap tier
Availability: The 256K context here is DeepInfra's, and it is what you get on this endpoint. NVIDIA's own model page and other catalogues advertise 1M tokens for the open weights, so the same model served elsewhere or self-hosted may accept a longer prompt than this row's price buys.

Today's price · per 1M tokens

Input

$0.080

per 1M tokens

Output

$0.200

per 1M tokens

Blended

$0.110

blended $/1M — 3:1 weighted input:output

Cached input

$0.040

50% of input — prompt caching

Source: official DeepInfra pricing · read Aug 17, 2026 MODELPRICEWATCH.COM · 2026-08-18

Price receipt

We read DeepInfra's own pricing page on Aug 17, 2026 and found NVIDIA-Nemotron-3.5-Lightning listed with a price on the same row — that read is where the number above comes from, recorded as a new model. We keep the page text we read; its content hash is 685265e497.

Overview

NVIDIA Nemotron 3.5 Lightning served on DeepInfra at $0.08/$0.20 per 1M tokens, with cached input at $0.04. The cheapest of the three hosts serving it (CoreWeave and Venice both list $0.10/$0.25). NVIDIA publishes no per-token price of its own for this model, so a host rate is the only price it has.

Run it yourself

Deploy this open model on rented GPUs

Open weights mean you can self-host instead of paying per-token API prices. These platforms let you serve it on demand — often cheaper at scale.

Partner links — we may earn a commission if you sign up, at no cost to you. We only list platforms we'd recommend regardless.

Capabilities

struck through = not supported
Input 1/5
Text ✓ Image Audio Video PDF
Output 1/5
Text ✓ Image Audio Video Embedding
Features 3/9
Prompt caching ✓ Reasoning Coding Fast inference Long context ✓ Open weights ✓ Multimodal Web search Realtime

Benchmark performance

accuracy % · higher is better
Percentile vs all tracked models 29.4th3 independent measurements
Percentile per $/Mtok 267.3

We hold no per-benchmark accuracy scores for this model yet, so it has no accuracy average. That is a gap in our coverage, not a sign the model is untested — independent evaluators often publish a composite index for a new model long before releasing its per-benchmark numbers. The percentile above is its standing across the independent composites below.

Independent composite scores
  • AA Intelligence Index: 23.6 (Artificial Analysis composite across reasoning, knowledge and coding evals)
  • AA Agentic Index: 13.8 (Tool use, planning, autonomy)
  • AA-Omniscience Index: -17.7 (−100–100)
How it stacks up
  • Ranks #106 of 161 comparably-measured models by percentile score, across 3 independent measurements
  • Ranks #5 of 161 comparably-tested models by normalized performance per dollar top value

Source: Artificial Analysis · updated Aug 17, 2026 · See full rankings →

Specifications

Provider
DeepInfra
Context window
256K tokens
Modality
text
Parameters
30B (3B active)
Open source
Yes — open weights available
Released
Aug 11, 2026
Status
Current
Last updated
Aug 17, 2026
Tags
open-weightscachingcheap

Availability verified: Aug 17, 2026 — listed on DeepInfra's own page