← Model list
llama-3.1-nemotron-ultra-253b-v1
nvidia·nvidia/llama-3.1-nemotron-ultra-253b-v1·GA·Closed·nemotron series
A reasoning-optimized LLM based on Llama 3.1, Nemotron Ultra 253B delivers strong performance in tasks like RAG and tool use, with high efficiency and reduced latency.
At a glance
Reference price$0.60 / $1.80 per 1M
Cache read $0.06 · Nebius Token Factory
Lowest paid$0.598 / $1.79
Cortecs Gateway · 1× spread
Context128,000
Output limit4,096
Capabilities
✓ Reasoning✓ Tool use✓ Structured outputTemperatureAttachments
⚠ Providers report capability flags inconsistently
Modalities
Text
Knowledge cutoff2024-12
Released / updated2025-04-07 / 2026-02-04
Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier Reasoning
Intelligence8.9
Value10
Coding—
Agentic—
Output speed48 tok/s
TTFT2.34 s
Value formulaIQ 8.9 ÷ min blended $0.897 = 10
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 2 providers2 with public prices
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Cortecs | Gateway | $0.598 | $1.79 | — | — | 128,000 | 128,000 | |
| Nebius Token Factory nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 | Cloud | $0.60 | $1.80 | $0.06 | $0.75 | 128,000 | 4,096 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1Nebius Token Factory$145.20
2Cortecs$209.30
The cheapest paid channel is the only channel.
Benchmark
No upstream benchmark data for this model. For quality, see the Artificial Analysis intelligence score above.
Reasoning control
Upstream provides no control info
1 / 2 providers expose no reasoning control (reasoning_options: []).
Related models
nemotron-lightning-3.5-30b-a3bsame series$0.045 / $0.18Nemotron 3.5 Lightning 30B A3Bsame series$0 / $0Nvidia Nemotron 3.5 Lightningsame series$0.05 / $0.20DeepSeek V4 Flashcheaper alternative$0.14 / $0.28MiniMax-M2.7cheaper alternative$0.30 / $1.20MiniMax-M3cheaper alternative$0.30 / $1.20GPT-5.6 Lunacheaper alternative$0.20 / $1.20
Price historyone sample accumulated per data sync
Input listOutput listMin blended