LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

nemotron-3-ultra-nvfp4

nvidia·nvidia/nemotron-3-ultra-nvfp4·GA·Closed·nemotron series
Nemotron-3-Ultra-550B-A55B-NVFP4 is a frontier-scale large language model (LLM) trained by NVIDIA, designed to deliver strong agentic, reasoning, and conversational capabilities. It is optimized for the most demanding workloads, including complex multi-step agents, long-context analysis, and high-accuracy reasoning over code, math, and science. The model employs a hybrid Latent Mixture-of-Experts (LatentMoE) architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. Like the Super model, the Ultra model incorporates Multi-Token Prediction (MTP) layers for faster text generation and improved quality, and it is trained using an NVFP4 pre-training recipe to maximize compute efficiency. The model has 55B active parameters and 550B parameters in total.

Specs & pricing

Input / output per 1M tokens
Reference price·Requesty
$0.60 / $2.40
Blended $1.05 · Cache read $0.12
Lowest paid·RequestyGateway
$0.60 / $2.40
Blended $1.05
Context
262,144
Output limit
262,144
Knowledge cutoff
—
Released / updated
2026-06-13 / 2026-06-13
Capabilities
✓ Reasoning✓ Tool useStructured output? TemperatureAttachments
Modalities
Text

Available at 1 providers1 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
RequestyGateway$0.60$2.40$0.12—262,144262,144

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

effort = noneeffort = loweffort = mediumeffort = higheffort = max

Your usage cost

1Requesty$182.40
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$2.40$02026-08-192026-08-282026-08-19 · Input list $0.602026-08-21 · Input list $0.542026-08-28 · Input list $0.602026-08-19 · Output list $2.402026-08-21 · Output list $2.162026-08-28 · Output list $2.402026-08-19 · Min blended $1.052026-08-21 · Min blended $0.952026-08-28 · Min blended $1.05

Related models

Nemotron 3.5 Lightning 30B A3Bsame series$0 / $0Nemotron 3.5 Lightning (free)same series$0 / $0Nemotron 3.5 Lightningsame series$0.07 / $0.20GLM-5.3-Flashcheaper alternative$0.15 / $0.50DeepSeek V4.1 Flashcheaper alternative$0.15 / $0.60GPT-5.6 Lunacheaper alternative$0.20 / $1.20GPT-5.4 nanocheaper alternative$0.20 / $1.25
Data partly from models.dev (MIT) · Requesty official docs ↗