← Model list
Qwen3.5 4B
alibaba·alibaba/qwen3.5-4b·GA·Open weights·qwen series·NEW
Qwen instruction model for multilingual chat and tool use
At a glance
Reference price$0.10 / $0.20 per 1M
Cache read $0.05 · NanoGPT
Lowest paid$0.04 / $0.07
EmpirioLabs AI Gateway · 2.5× spread
1 more $0 channels
Context262,144
⚠ Providers report 32,768–262,144; the table below is authoritative
Output limit32,768
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ Temperature✓ Attachments
⚠ Providers report capability flags inconsistently
Modalities
TextImageVideo
Knowledge cutoff—
Released / updated2025-11-01 / 2026-08-16
Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier Reasoning
Intelligence20.4
Value429
Coding22.6
Agentic—
Output speed30 tok/s
TTFT0.71 s
Value formulaIQ 20.4 ÷ min blended $0.048 = 429
Reasoning tier → intelligence / speed (higher tier = stronger but slower)
Non-reasoningIQ 16.1 · 28 tok/s
ReasoningIQ 20.4 · 30 tok/s
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 3 providers2 with public prices · 1 free
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| QVAC | Gateway | Free | — | — | 32,768 ⚠ | 8,192 | ||
| EmpirioLabs AI qwen3-5-4b | Gateway | $0.04 | $0.07 | $0.02 | — | 262,144 | 32,768 | |
| NanoGPT | Gateway | $0.10 | $0.20 | $0.05 | — | 262,144 | 32,768 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1EmpirioLabs AI$9.10
2NanoGPT$24.00
1 more channels offer $0 (QVAC); free tiers usually have rate limits and no SLA, excluded from ranking.
The cheapest paid channel is the only channel.
Benchmark
No upstream benchmark data for this model. For quality, see the Artificial Analysis intelligence score above.
Reasoning control
Toggle (on / off)effort = noneeffort = loweffort = mediumeffort = higheffort = maxbudget_tokens ≥ 1,024 ≤ 32,768
1 / 3 providers expose no reasoning control (reasoning_options: []).
Related models
Qwen3.5 0.8Bsame series$0.06 / $0.12Qwen3.8 27B TEEsame series$0.40 / $3.00Qwen3.8 27Bsame series$0.575 / $3.45GLM-4.7-Flashcheaper alternative$0 / $0Hy3cheaper alternative$0 / $0Nemotron 3 Nano 30B A3Bcheaper alternative$0 / $0Qwen3.7 Flashcheaper alternative$0.03 / $0.118
Price historyone sample accumulated per data sync
Input listOutput listMin blended