Inkling 256K is the extended context variant of Inkling, a large MoE hybrid reasoning model from Thinking Machines with audio and vision input support and a 256K context window.
inkling-256k by thinkingmachines is offered by 2 providers on this page. Its official list price is $3.74 per 1M input tokens and $9.36 per 1M output tokens. The lowest paid channel is Requesty at $1.87 / $4.68 per 1M, about 2.0× below the list price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it.
The context window is 262,144 tokens, with an output limit of 262,144 tokens. It supports reasoning and tool use. Accepted input modalities are Text and Image. Providers report the capability flags inconsistently, so verify a specific feature against the host you plan to use.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Requesty | Gateway | $1.87 | $4.68 | $0.374 | — | 262,144 | 32,768 | |
| Thinking MachinesOfficial thinkingmachines/Inkling:peft:262144 | First-party | $3.74 | $9.36 | $0.748 | — | 262,144 | 262,144 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.