A reasoning-optimized LLM based on Llama 3.1, Nemotron Ultra 253B delivers strong performance in tasks like RAG and tool use, with high efficiency and reduced latency.
llama-3.1-nemotron-ultra-253b-v1 by misc is currently listed from a single provider. Its reference price is $0.598 per 1M input tokens and $1.79 per 1M output tokens.
The context window is 128,000 tokens, with an output limit of 128,000 tokens. It supports reasoning, tool use, and structured output.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Cortecs | Gateway | $0.598 | $1.79 | — | — | 128,000 | 128,000 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
1 / 1 providers expose no reasoning control (reasoning_options: []).
Price history accumulates from each data sync; currently only 1 sample(s) (2026-09-03). Each future sync adds a point, and once accumulated a line is drawn here.