pixtral-12b-2409
Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.
Quality & performance
Artificial Analysis doesn't cover this model (267 of 2059 have data). Quality data comes from independent evals covering widely used models.
Available at 2 providers2 with public prices
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
Benchmark
No upstream benchmark data for this model. For quality, see the Artificial Analysis intelligence score above.
Reasoning control
1 / 2 providers expose no reasoning control (reasoning_options: []).