Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.
pixtral-12b-2409 by misc is currently listed from a single provider. Its reference price is $0.223 per 1M tokens.
The context window is 128,000 tokens, with an output limit of 4,096 tokens. It supports reasoning, tool use, and structured output. Accepted input modalities are Text and Image.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Cortecs | Gateway | $0.22 | $0.22 | — | — | 128,000 | 4,096 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
1 / 1 providers expose no reasoning control (reasoning_options: []).
Price history accumulates from each data sync; currently only 1 sample(s) (2026-10-11). Each future sync adds a point, and once accumulated a line is drawn here.