LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

inkling-256k

thinkingmachines·thinkingmachines/inkling-256k·GA·Closed·ling series·Video gen·NEW
Inkling 256K is the extended context variant of Inkling, a large MoE hybrid reasoning model from Thinking Machines with audio and vision input support and a 256K context window.

inkling-256k by thinkingmachines is offered by 2 providers on this page. Its official list price is $3.74 per 1M input tokens and $9.36 per 1M output tokens. The lowest paid channel is Requesty at $1.87 / $4.68 per 1M, about 2.0× below the list price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it.

The context window is 262,144 tokens, with an output limit of 262,144 tokens. It supports reasoning and tool use. Accepted input modalities are Text and Image. Providers report the capability flags inconsistently, so verify a specific feature against the host you plan to use.

Specs & pricing

Input / output per 1M tokens
Official price·Thinking Machines
$3.74 / $9.36
Blended $5.14 · Cache read $0.748
Lowest paid·RequestyGateway
$1.87 / $4.68
Blended $2.57 · 2× spread
Context
262,144
Output limit
262,144
Knowledge cutoff
—
Released / updated
2026-07-16 / 2026-07-16
Capabilities
✓ Reasoning✓ Tool useStructured output✓ Temperature✓ Attachments
⚠ Providers report capability flags inconsistently
Modalities
TextImage

Available at 2 providers2 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
RequestyGateway$1.87$4.68$0.374—262,14432,768
Thinking MachinesOfficial
thinkingmachines/Inkling:peft:262144
First-party$3.74$9.36$0.748—262,144262,144

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1Requesty$428.48
2Thinking Machines · Official$856.96
Switch to Requesty to save $428.48/mo (50%)
Note: this is a gateway; verify availability and rate limits yourself.

Reasoning control

effort = noneeffort = loweffort = mediumeffort = higheffort = maxbudget_tokensToggle (on / off)effort = xhigh

Related models

inLing 3.0 Flash Sante (Free)same series$0 / $0inLing 3.0 Flash Santesame series$0 / $0inLing 3.0 Flash Finsame series$0.06 / $0.18GLM-5.2cheaper alternative$1.40 / $4.40DeepSeek V4 Flashcheaper alternative$0.14 / $0.28DeepSeek V4 Procheaper alternative$0.435 / $0.87Kimi K2.7 Codecheaper alternative$0.95 / $4.00

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$9.36$02026-08-192026-08-282026-08-19 · Input list $3.742026-08-21 · Input list $3.742026-08-28 · Input list $3.742026-08-19 · Output list $9.362026-08-21 · Output list $9.362026-08-28 · Output list $9.362026-08-19 · Min blended $2.572026-08-21 · Min blended $2.062026-08-28 · Min blended $2.57
Data partly from models.dev (MIT) · Thinking Machines official docs ↗