← Model list
Gemini 3.5 Flash
google·google/gemini-3.5-flash·GA·Closed·gemini-flash series
Fast Gemini model balancing multimodal reasoning, tool use, and cost
At a glance
Official price$1.50 / $9.00 per 1M
Cache read $0.15 · Vertex
Lowest paid$0.186 / $1.11
UnoRouter Gateway · 8.9× spread
Context1,048,576
⚠ Providers report 200,000–1,048,576; the table below is authoritative
Output limit65,536
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ Temperature✓ Attachments
Modalities
TextImageAudioVideoPDF
Knowledge cutoff2025-01
Released / updated2026-05-19 / 2026-05-19
Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier high
Intelligence52
Value124
Coding70.1
Agentic39.7
Output speed171 tok/s
TTFT27.29 s
Cost per task$0.6934
Value formulaIQ 52 ÷ min blended $0.418 = 124
Reasoning tier → intelligence / speed (higher tier = stronger but slower)
minimalIQ 35.8 · 154 tok/s
mediumIQ 46.7 · 180 tok/s
highIQ 52 · 171 tok/s
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 29 providers29 with public prices
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| UnoRouter | Gateway | $0.186 | $1.11 | — | — | 1,048,576 | 65,536 | |
| Requesty | Gateway | $1.35 | $8.10 | $0.135 | $1.42 | 1,048,576 | 65,535 | |
| NanoGPT | Gateway | $1.50 | $9.00 | $0.15 | $0.083 | 1,048,576 | 65,536 | |
| VertexOfficial | First-party | $1.50 | $9.00 | $0.15 | — | 1,048,576 | 65,536 | |
| GoogleOfficial | First-party | $1.50 | $9.00 | $0.15 | — | 1,048,576 | 65,536 | |
| Impossibl | Gateway | $1.50 | $9.00 | $0.15 | — | 1,048,576 | 65,536 | |
| OpenRouter | Gateway | $1.50 | $9.00 | $0.15 | $0.083 | 1,048,576 | 65,536 | |
| CrossModel | Gateway | $1.50 | $9.00 | $0.15 | $1.50 | 1,048,576 | 65,536 | |
| NEAR AI Cloud | Gateway | $1.50 | $9.00 | $0.15 | — | 1,048,576 | 65,536 | |
| Merge Gateway | Gateway | $1.50 | $9.00 | $0.15 | — | 1,048,576 | 65,536 | |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1UnoRouter$92.85
2Requesty$529.20
3NanoGPT$588.00
4Vertex · Official$588.00
5Google · Official$588.00
6Impossibl$588.00
Switch to UnoRouter to save $495.15/mo (84%)
Note: this is a gateway; verify availability and rate limits yourself.
Benchmark10 items
| Name | Conditions | Score | Metric | Source |
|---|---|---|---|---|
| Terminal-Bench | harness: Terminus-2 · v2.1 | 76.2 | success rate | Source ↗ |
| SWE-Bench Pro | variant: single attempt · dataset: public | 55.1 | resolve rate | Source ↗ |
| MCP Atlas | — | 83.6 | success rate | Source ↗ |
| Toolathlon | — | 56.5 | success rate | Source ↗ |
| OSWorld-Verified | — | 78.4 | success rate | Source ↗ |
| MMMU Pro | variant: no tools | 83.6 | accuracy | Source ↗ |
| CharXiv Reasoning | variant: no tools | 84.2 | accuracy | Source ↗ |
| Humanity's Last Exam | dataset: full set, text + MM | 40.2 | accuracy | Source ↗ |
| ARC-AGI-2 | — | 72.1 | accuracy | Source ↗ |
| GDPval-AA | — | 1656 | Elo | Source ↗ |
The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.
Reasoning control
effort = noneeffort = loweffort = mediumeffort = higheffort = maxbudget_tokens ≥ 256 ≤ 24,000effort = minimalToggle (on / off)effort = xhigh
3 / 29 providers expose no reasoning control (reasoning_options: []).
Related models
Gemini 3.7 Flashsame series$0.75 / $3.75Gemini 3.7 Flash (Google Vertex AI)same series$0.75 / $3.75Gemini 3.7 Flash (Google AI Studio)same series$0.75 / $3.75DeepSeek V4 Procheaper alternative$0.435 / $0.87DeepSeek V4 Flashcheaper alternative$0.14 / $0.28DeepSeek V4 Flash 0731cheaper alternative$0.479 / $1.44MiniMax-M3cheaper alternative$0.30 / $1.20
Price historyone sample accumulated per data sync
Input listOutput listMin blended