LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

Nemotron 3.5 Lightning

nvidia·nvidia/nvidia-nemotron-3.5-lightning-30b-a3b·GA·Open weights·nemotron series·NEW
Nemotron 3.5 Lightning is an MoE model built for fast, reliable agentic tasks across use cases such as financial services, cybersecurity, telecom, and retail.

Nemotron 3.5 Lightning by nvidia is currently listed from a single provider. Its reference price is $0.07 per 1M input tokens and $0.20 per 1M output tokens.

Artificial Analysis rates it 12.9 on the Intelligence Index, with 26.8 for coding and 3.5 for agentic tasks. Against its lowest blended price of $0.102 per 1M, that is roughly 126 index points per dollar, which is the value ratio the leaderboards rank on. Median output speed is 299 tokens per second, with 0.56s to the first token. Running one task of the Artificial Analysis suite costs about $0.0931, which reflects how many tokens its reasoning consumes rather than the unit price alone.

The context window is 262,144 tokens, with an output limit of 262,144 tokens. It supports reasoning, tool use, and structured output. The weights are open, so it can also be self-hosted or served through a gateway of your choice.

Specs & pricing

Input / output per 1M tokens
Reference price·CoreWeave
$0.07 / $0.20
Blended $0.10 · Cache read $0.04
Lowest paid·CoreWeaveCloud
$0.07 / $0.20
Blended $0.10
Context
262,144
Output limit
262,144
Knowledge cutoff
—
Released / updated
2026-08-11 / 2026-08-11
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
Modalities
Text

Quality & performanceArtificial Analysis · Intelligence Index v4.3

Intelligence12.9
Value126
Coding26.8
Agentic3.5
Output speed299 tok/s
TTFT0.56 s
Cost per task$0.093
Value formulaIQ 12.9 ÷ min blended $0.102 = 126

Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.

Available at 1 providers1 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
CoreWeave
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B
Cloud$0.07$0.20$0.04—262,144262,144

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

Toggle (on / off)

Your usage cost

1CoreWeave$20.40
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.25$02026-08-112026-09-152026-08-11 · Input list $0.102026-08-12 · Input list $0.102026-08-13 · Input list $0.102026-09-15 · Input list $0.072026-08-11 · Output list $0.252026-08-12 · Output list $0.252026-08-13 · Output list $0.252026-09-15 · Output list $0.202026-08-11 · Min blended $0.142026-08-12 · Min blended $0.142026-08-13 · Min blended $0.142026-09-15 · Min blended $0.10

Artificial Analysis evaluations6 items

GPQA Diamond74.3%
Humanity's Last Exam10.6%
SciCode32.1%
Terminal-Bench 2.124.3%
τ³-Bench Banking8.9%
AA-LCR60.3%

Individual evaluations run by Artificial Analysis, on the same reasoning tier as the intelligence score above. Each benchmark has its own task set and harness, so rows are not comparable with one another. The Intelligence Index above draws on a different, newer set of evaluations.

Related models

Nemotron 3.5 Lightning 30B A3Bsame series$0 / $0Nemotron 3.5 Lightning (free)same series$0 / $0Nvidia Nemotron 3.5 Lightning Thinkingsame series$0.05 / $0.20Hy3cheaper alternative$0 / $0Qwen3.7 Flashcheaper alternative$0.03 / $0.13Laguna S 2.1cheaper alternative$0 / $0
Data partly from models.dev (MIT) · CoreWeave official docs ↗