LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list
mi

llama-3.1-nemotron-ultra-253b-v1

misc·misc/llama-3.1-nemotron-ultra-253b-v1·GA·Closed
A reasoning-optimized LLM based on Llama 3.1, Nemotron Ultra 253B delivers strong performance in tasks like RAG and tool use, with high efficiency and reduced latency.

llama-3.1-nemotron-ultra-253b-v1 by misc is currently listed from a single provider. Its reference price is $0.598 per 1M input tokens and $1.79 per 1M output tokens.

The context window is 128,000 tokens, with an output limit of 128,000 tokens. It supports reasoning, tool use, and structured output.

Specs & pricing

Input / output per 1M tokens
Reference price·Cortecs
$0.598 / $1.79
Blended $0.897 · Cache read —
Lowest paid·CortecsGateway
$0.598 / $1.79
Blended $0.897
Context
128,000
Output limit
128,000
Knowledge cutoff
—
Released / updated
2025-04-07 / 2025-04-07
Capabilities
✓ Reasoning✓ Tool use✓ Structured outputTemperatureAttachments
Modalities
Text

Available at 1 providers1 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
CortecsGateway$0.598$1.79——128,000128,000

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1Cortecs$209.30
The cheapest paid channel is the only channel.

Reasoning control

Upstream provides no control info

1 / 1 providers expose no reasoning control (reasoning_options: []).

Related models

DeepSeek V4 Flashcheaper alternative$0.14 / $0.28GLM-5.3-Flashcheaper alternative$0.075 / $0.25MiniMax-M2.7cheaper alternative$0.30 / $1.20MiniMax-M3cheaper alternative$0.30 / $1.20

Price historyone sample accumulated per data sync

Input list $0.598Output list $1.79Min blended $0.897

Price history accumulates from each data sync; currently only 1 sample(s) (2026-09-03). Each future sync adds a point, and once accumulated a line is drawn here.

Data partly from models.dev (MIT) · Cortecs official docs ↗