← Model list
Nemotron 3 Ultra 550B A55B
nvidia·nvidia/nemotron-3-ultra-550b-a55b·GA·Open weights·nemotron series·NEW
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
At a glance
Official price$0.50 / $2.50 per 1M
Cache read $0.15 · Nvidia
Lowest paid$0.10 / $0.10
routing.run Gateway · 10× spread
3 more $0 channels
Context1,000,000
⚠ Providers report 131,072–1,048,576; the table below is authoritative
Output limit65,536
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
Modalities
Text
Knowledge cutoff—
Released / updated2026-06-04 / 2026-06-04
Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier Reasoning
Intelligence38.3
Value383
Coding49.3
Agentic27.5
Output speed132 tok/s
TTFT1.98 s
Cost per task$0.3827
Value formulaIQ 38.3 ÷ min blended $0.100 = 383
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 14 providers11 with public prices · 3 free
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Kenari | Gateway | Free | — | — | 1,000,000 | 128,000 | ||
| UnoRouter nemotron-3-ultra-550b-a55b:free | Gateway | Free | — | — | 1,000,000 | 128,000 | ||
| Requesty | Gateway | Free | — | — | 1,048,576 ⚠ | 65,536 | ||
| routing.run nemotron-3-ultra | Gateway | $0.10 | $0.10 | — | — | 131,072 ⚠ | 32,000 | |
| Eden AI | Gateway | $0.50 | $2.20 | $0.10 | — | 262,144 ⚠ | 128,000 | |
| DevPass (LLM Gateway) nemotron-3-ultra-550b | Gateway | $0.50 | $2.20 | $0.10 | — | 1,048,576 ⚠ | 128,000 | |
| Kilo Gateway | Gateway | $0.50 | $2.20 | $0.10 | — | 512,288 ⚠ | 512,288 | |
| NvidiaOfficial | First-party | $0.50 | $2.50 | $0.15 | — | 1,000,000 | 65,536 | |
| Pioneer nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 | Gateway | $0.50 | $2.50 | $0.15 | $0.50 | 1,000,000 | 65,000 | |
| Fireworks AI accounts/fireworks/models/nemotron-3-ultra-nvfp4 | Cloud | $0.60 | $2.40 | $0.119 | — | 262,144 ⚠ | 128,000 | |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1routing.run$25.00
2Eden AI$162.00
3DevPass (LLM Gateway)$162.00
4Kilo Gateway$162.00
5Fireworks AI$182.28
6Nvidia · Official$183.00
3 more channels offer $0 (Kenari, UnoRouter, Requesty); free tiers usually have rate limits and no SLA, excluded from ranking.
Switch to routing.run to save $158.00/mo (86%)
Note: this is a gateway; verify availability and rate limits yourself.
Benchmark11 items
| Name | Conditions | Score | Metric | Source |
|---|---|---|---|---|
| SWE-Bench Verified | — | 70.7 | resolved | Source ↗ |
| SWE-Bench Multilingual | — | 67.7 | resolve rate | Source ↗ |
| Terminal-Bench | v2.1 | 56.4 | success rate | Source ↗ |
| GPQA | variant: no tools | 87 | accuracy | Source ↗ |
| Humanity's Last Exam | variant: no tools | 26.7 | accuracy | Source ↗ |
| Humanity's Last Exam | variant: with tools | 37.4 | accuracy | Source ↗ |
| LiveCodeBench | vv6 | 89 | pass@1 | Source ↗ |
| MMLU-Pro | — | 86.8 | accuracy | Source ↗ |
| BrowseComp | — | 44.4 | accuracy | Source ↗ |
| IFBench | variant: prompt loose | 81.7 | accuracy | Source ↗ |
| GDPval | — | 46.7 | wins or ties | Source ↗ |
The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.
Reasoning control
effort = mediumeffort = higheffort = noneeffort = loweffort = maxbudget_tokensToggle (on / off)
2 / 14 providers expose no reasoning control (reasoning_options: []).
Related models
nemotron-lightning-3.5-30b-a3bsame series$0.045 / $0.18Nemotron 3.5 Lightning 30B A3Bsame series$0 / $0Nvidia Nemotron 3.5 Lightningsame series$0.05 / $0.20DeepSeek V4 Procheaper alternative$0.435 / $0.87DeepSeek V4 Flashcheaper alternative$0.14 / $0.28MiniMax-M3cheaper alternative$0.30 / $1.20GPT-5.6 Lunacheaper alternative$0.20 / $1.20
Price historyone sample accumulated per data sync
Input listOutput listMin blended