← Model list

nvidia-nemotron-3-ultra

nvidia·nvidia/nvidia-nemotron-3-ultra·GA·Closed·nemotron series·NEW

NVIDIA Nemotron 3 Ultra is NVIDIA's strongest open-weights reasoning model, positioned near GPT-5.4 Mini (xhigh) and ahead of DeepSeek V4-Flash and Qwen3.5-397B-A17B.

At a glance
Reference price$0.625 / $3.13 per 1M
Cache read $0.188 · Venice AI
Lowest paid$0.45 / $2.25
Requesty Gateway · 1.4× spread
Context256,000
⚠ Providers report 256,000–262,144; the table below is authoritative
Output limit32,768
Capabilities
ReasoningTool useStructured outputTemperatureAttachments
⚠ Providers report capability flags inconsistently
Modalities
Text
Knowledge cutoff
Released / updated2026-06-23 / 2026-06-23

Quality & performance

Artificial Analysis doesn't cover this model (267 of 2059 have data). Quality data comes from independent evals covering widely used models.

Available at 2 providers2 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
RequestyGateway$0.45$2.25262,144131,072
Venice AI
nvidia-nemotron-3-ultra-550b-a55b
Gateway$0.625$3.13$0.188256,00032,768

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1Requesty$202.50
2Venice AI$228.75
The cheapest paid channel is the only channel.

Benchmark

No upstream benchmark data for this model. For quality, see the Artificial Analysis intelligence score above.

Reasoning control

effort = noneeffort = loweffort = mediumeffort = higheffort = maxbudget_tokens

Related models

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$3.13$02026-08-192026-08-212026-08-19 · Input list $0.6252026-08-21 · Input list $0.6252026-08-19 · Output list $3.132026-08-21 · Output list $3.132026-08-19 · Min blended $1.002026-08-21 · Min blended $0.90