Fast Gemini model balancing multimodal reasoning, tool use, and cost
Gemini 3 Flash by google is offered by 3 providers on this page. Its reference price is $0.50 per 1M input tokens and $3.00 per 1M output tokens. The lowest paid channel is Poe at $0.40 / $2.40 per 1M, about 1.3× below the reference price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it.
The context window is 1,048,576 tokens at the reference host, but hosts report different limits, from 1,000,000 to 1,048,576, so the usable window depends on the provider you pick. It supports reasoning, tool use, and structured output. Accepted input modalities are Text, Image, Video, Audio, and PDF. Its training knowledge cuts off at 2025-03.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Poe | Gateway | $0.40 | $2.40 | $0.04 | — | 1,048,576 | 65,536 | |
| Vercel AI Gateway | Cloud | $0.50 | $3.00 | $0.05 | — | 1,000,000 ⚠ | 65,000 | |
| OpenCode Zen | Gateway | $0.50 | $3.00 | $0.05 | — | 1,048,576 | 65,536 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.