LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

Nemotron 3 Ultra

nvidia·nvidia/nvidia-nemotron-3-ultra-550b-a55b·GA·Open weights·nemotron series
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

Specs & pricing

Input / output per 1M tokens
Reference price·Venice AI
$0.63 / $3.13
Blended $1.25 · Cache read $0.19
Lowest paid·Ollama CloudCloud
$0.10 / $3.00
Blended $0.83 · 9× spread
Context
256,000
Output limit
32,768
Knowledge cutoff
—
Released / updated
2026-06-04 / 2026-06-23
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
⚠ Providers report capability flags inconsistently
Modalities
Text

Available at 7 providers7 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
Ollama Cloud
nemotron-3-ultra
Cloud$0.10$3.00$0.10—262,144 ⚠128,000
CoreWeave
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
Cloud$0.50$2.15$0.10—262,144 ⚠262,144
Requesty
nvidia-nemotron-3-ultra
Gateway$0.50$2.50——262,144 ⚠131,072
Vercel AI Gateway
nvidia/nemotron-3-ultra-550b-a55b
Cloud$0.60$2.40$0.12—1,000,000 ⚠65,000
Baseten
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B
Cloud$0.60$2.40$0.12—202,800 ⚠202,800
DigitalOcean
nemotron-3-ultra-550b
Cloud$0.90$1.70$0.18—131,072 ⚠131,072
Venice AIGateway$0.63$3.13$0.19—256,00032,768

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

Toggle (on / off)effort = noneeffort = loweffort = mediumeffort = higheffort = max

Interleaved thinking (reasoning between tool calls) is declared by 1 of 7 providers.

Your usage cost

1CoreWeave$159.50
2Ollama Cloud$170.00
3DigitalOcean$178.60
4Vercel AI Gateway$182.40
5Baseten$182.40
6Requesty$225.00
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$3.13$02026-08-052026-09-242026-08-05 · Input list $0.752026-08-06 · Input list $0.752026-08-07 · Input list $0.752026-08-08 · Input list $0.752026-08-09 · Input list $0.752026-08-10 · Input list $0.752026-08-11 · Input list $0.752026-08-12 · Input list $0.752026-08-13 · Input list $0.602026-09-24 · Input list $0.632026-08-05 · Output list $2.752026-08-06 · Output list $2.752026-08-07 · Output list $2.752026-08-08 · Output list $2.752026-08-09 · Output list $2.752026-08-10 · Output list $2.752026-08-11 · Output list $2.752026-08-12 · Output list $2.752026-08-13 · Output list $2.402026-09-24 · Output list $3.132026-08-05 · Min blended $1.052026-08-06 · Min blended $1.052026-08-07 · Min blended $1.052026-08-08 · Min blended $1.052026-08-09 · Min blended $1.052026-08-10 · Min blended $1.052026-08-11 · Min blended $1.052026-08-12 · Min blended $1.052026-08-13 · Min blended $1.052026-09-24 · Min blended $0.83

Related models

Nemotron 3.5 Lightning 30B A3Bsame series$0 / $0Nemotron 3.5 Lightning (free)same series$0 / $0Nemotron 3.5 Lightningsame series$0.07 / $0.20GLM-5.3-Flashcheaper alternative$0.15 / $0.50DeepSeek V4.1 Flashcheaper alternative$0.15 / $0.60DeepSeek V4 Flash 0731cheaper alternative$0.45 / $1.34GPT-5.6 Lunacheaper alternative$0.20 / $1.20
Data partly from models.dev (MIT) · Venice AI official docs ↗