Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Mercury 2.5 Preview by inception is offered by 3 providers on this page. Its reference price is $0.20 per 1M input tokens and $0.75 per 1M output tokens. The lowest paid channel is NanoGPT at $0.04 / $0.15 per 1M, about 5.0× below the reference price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it.
The context window is 260,000 tokens, with an output limit of 65,536 tokens. It supports reasoning, tool use, and structured output.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| NanoGPT | Gateway | $0.04 | $0.15 | $0.004 | — | 260,000 | 65,536 | |
| OpenRouter | Gateway | $0.04 | $0.15 | $0.004 | — | 260,000 | 65,536 | |
| Kilo Gateway | Gateway | $0.20 | $0.75 | $0.02 | — | 260,000 | 65,536 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Price history accumulates from each data sync; currently only 1 sample(s) (2026-09-02). Each future sync adds a point, and once accumulated a line is drawn here.