← Model list

Nemotron 3 Ultra 550B A55B

nvidia·nvidia/nemotron-3-ultra-550b-a55b·GA·Open weights·nemotron series·NEW

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

At a glance
Official price$0.50 / $2.50 per 1M
Cache read $0.15 · Nvidia
Lowest paid$0.10 / $0.10
routing.run Gateway · 10× spread
3 more $0 channels
Context1,000,000
⚠ Providers report 131,072–1,048,576; the table below is authoritative
Output limit65,536
Capabilities
ReasoningTool useStructured outputTemperatureAttachments
Modalities
Text
Knowledge cutoff
Released / updated2026-06-04 / 2026-06-04

Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier Reasoning

Intelligence38.3
Value383
Coding49.3
Agentic27.5
Output speed132 tok/s
TTFT1.98 s
Cost per task$0.3827
Value formulaIQ 38.3 ÷ min blended $0.100 = 383

Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.

Available at 14 providers11 with public prices · 3 free

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
KenariGatewayFree1,000,000128,000
UnoRouter
nemotron-3-ultra-550b-a55b:free
GatewayFree1,000,000128,000
RequestyGatewayFree1,048,57665,536
routing.run
nemotron-3-ultra
Gateway$0.10$0.10131,07232,000
Eden AIGateway$0.50$2.20$0.10262,144128,000
DevPass (LLM Gateway)
nemotron-3-ultra-550b
Gateway$0.50$2.20$0.101,048,576128,000
Kilo GatewayGateway$0.50$2.20$0.10512,288512,288
NvidiaOfficialFirst-party$0.50$2.50$0.151,000,00065,536
Pioneer
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16
Gateway$0.50$2.50$0.15$0.501,000,00065,000
Fireworks AI
accounts/fireworks/models/nemotron-3-ultra-nvfp4
Cloud$0.60$2.40$0.119262,144128,000

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1routing.run$25.00
2Eden AI$162.00
3DevPass (LLM Gateway)$162.00
4Kilo Gateway$162.00
5Fireworks AI$182.28
6Nvidia · Official$183.00

3 more channels offer $0 (Kenari, UnoRouter, Requesty); free tiers usually have rate limits and no SLA, excluded from ranking.

Switch to routing.run to save $158.00/mo (86%)
Note: this is a gateway; verify availability and rate limits yourself.

Benchmark11 items

NameConditionsScoreMetricSource
SWE-Bench Verified70.7resolvedSource ↗
SWE-Bench Multilingual67.7resolve rateSource ↗
Terminal-Benchv2.156.4success rateSource ↗
GPQAvariant: no tools87accuracySource ↗
Humanity's Last Examvariant: no tools26.7accuracySource ↗
Humanity's Last Examvariant: with tools37.4accuracySource ↗
LiveCodeBenchvv689pass@1Source ↗
MMLU-Pro86.8accuracySource ↗
BrowseComp44.4accuracySource ↗
IFBenchvariant: prompt loose81.7accuracySource ↗
GDPval46.7wins or tiesSource ↗

The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.

Reasoning control

effort = mediumeffort = higheffort = noneeffort = loweffort = maxbudget_tokensToggle (on / off)

2 / 14 providers expose no reasoning control (reasoning_options: []).

Related models

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$2.50$02026-08-052026-08-132026-08-05 · Input list $0.502026-08-06 · Input list $0.502026-08-07 · Input list $0.502026-08-08 · Input list $0.502026-08-09 · Input list $0.502026-08-10 · Input list $0.502026-08-11 · Input list $0.502026-08-12 · Input list $0.502026-08-13 · Input list $0.502026-08-05 · Output list $2.502026-08-06 · Output list $2.502026-08-07 · Output list $2.502026-08-08 · Output list $2.502026-08-09 · Output list $2.502026-08-10 · Output list $2.502026-08-11 · Output list $2.502026-08-12 · Output list $2.502026-08-13 · Output list $2.502026-08-05 · Min blended $0.102026-08-06 · Min blended $0.102026-08-07 · Min blended $0.102026-08-08 · Min blended $0.102026-08-09 · Min blended $0.102026-08-10 · Min blended $0.102026-08-11 · Min blended $0.102026-08-12 · Min blended $0.102026-08-13 · Min blended $0.10