LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

Nemotron Nano 12B v2 VL

nvidia·nvidia/nemotron-nano-12b-v2-vl·GA·Open weights·nemotron series
Nemotron multimodal model for visual reasoning and agentic AI workflows

Nemotron Nano 12B v2 VL by nvidia is offered by 3 providers on this page. Public prices are shown for 2 of them. Its official list price is $0 per 1M tokens. It also has 1 free ($0) channel; free tiers usually carry rate limits, and subscription-covered access bills $0 per token only after the subscription fee.

The context window is 128,000 tokens at the reference host, but hosts report different limits, from 128,000 to 131,072, so the usable window depends on the provider you pick. It supports reasoning and tool use. Accepted input modalities are Text, Image, and Video. The weights are open, so it can also be self-hosted or served through a gateway of your choice.

Specs & pricing

Input / output per 1M tokens
Official price·Nvidia
$0 / $0
Blended $0 · Cache read —
Lowest paid·DigitalOceanCloud
$0.20 / $0.60
Blended $0.30 · 1× spread
1 more $0 channels
Context
128,000
Output limit
128,000
Knowledge cutoff
—
Released / updated
2025-10-28 / 2025-10-28
Capabilities
✓ Reasoning✓ Tool use? Structured output✓ Temperature✓ Attachments
Modalities
TextImageVideo

Available at 3 providers2 with public prices · 1 free

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
NvidiaOfficialFirst-partyFree——128,000128,000
DigitalOceanCloud$0.20$0.60——128,00016,384
Vercel AI GatewayCloud$0.20$0.60——131,072 ⚠131,072

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1DigitalOcean$70.00
2Vercel AI Gateway$70.00

1 more channels offer $0 (Nvidia); free tiers usually have rate limits and no SLA, excluded from ranking.

The cheapest paid channel is the only channel.

Reasoning control

effort = noneeffort = loweffort = mediumeffort = higheffort = maxToggle (on / off)

1 / 3 providers expose no reasoning control (reasoning_options: []).

Related models

nemotron-lightning-3.5-30b-a3bsame series$0.05 / $0.20Nemotron 3.5 Lightning 30B A3Bsame series$0 / $0Nemotron 3.5 Lightning Freesame series$0 / $0

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.30$02026-08-052026-08-132026-08-05 · Input list $02026-08-06 · Input list $02026-08-07 · Input list $02026-08-08 · Input list $02026-08-09 · Input list $02026-08-10 · Input list $02026-08-11 · Input list $02026-08-12 · Input list $02026-08-13 · Input list $02026-08-05 · Output list $02026-08-06 · Output list $02026-08-07 · Output list $02026-08-08 · Output list $02026-08-09 · Output list $02026-08-10 · Output list $02026-08-11 · Output list $02026-08-12 · Output list $02026-08-13 · Output list $02026-08-05 · Min blended $0.302026-08-06 · Min blended $0.302026-08-07 · Min blended $0.302026-08-08 · Min blended $0.302026-08-09 · Min blended $0.302026-08-10 · Min blended $0.302026-08-11 · Min blended $0.302026-08-12 · Min blended $0.302026-08-13 · Min blended $0.30
Data partly from models.dev (MIT) · Nvidia official docs ↗