← Model list

Qwen3 8B

alibaba·alibaba/qwen3-8b·GA·Open weights·qwen series

Qwen instruction model for multilingual chat, reasoning, and tool use

At a glance
Official price$0.072 / $0.287 per 1M
Cache read · Alibaba (China)
Lowest paid$0.035 / $0.138
NovitaAI Gateway · 13.4× spread
Context131,072
⚠ Providers report 40,960–131,072; the table below is authoritative
Output limit8,192
Capabilities
ReasoningTool useStructured outputTemperatureAttachments
⚠ Providers report capability flags inconsistently
Modalities
Text
Knowledge cutoff2025-03-31
Released / updated2025-04 / 2025-04-29

Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier Reasoning

Intelligence8.3
Value137
Coding9
Agentic1.6
Output speed39 tok/s
TTFT3.76 s
Value formulaIQ 8.3 ÷ min blended $0.061 = 137
Reasoning tier → intelligence / speed (higher tier = stronger but slower)
Non-reasoningIQ 4.8 · 39 tok/s
ReasoningIQ 8.3 · 39 tok/s

Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.

Available at 6 providers6 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
NovitaAI
qwen/qwen3-8b-fp8
Gateway$0.035$0.138128,00020,000
Alibaba (China)OfficialFirst-party$0.072$0.287131,0728,192
Pioneer
Qwen/Qwen3-8B
Gateway$0.20$0.20$0.20$0.2040,96040,960
OpenRouterGateway$0.117$0.455131,0728,192
AlibabaOfficialFirst-party$0.18$0.70131,0728,192
NanoGPT
qwen/Qwen3-8B
Gateway$0.47$0.47$0.23541,00032,768

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1NovitaAI$13.90
2Alibaba (China) · Official$28.75
3OpenRouter$46.15
4Pioneer$50.00
5Alibaba · Official$71.00
6NanoGPT$89.30
Switch to NovitaAI to save $14.85/mo (52%)
Note: this is a gateway; verify availability and rate limits yourself.

Benchmark

No upstream benchmark data for this model. For quality, see the Artificial Analysis intelligence score above.

Reasoning control

Toggle (on / off)budget_tokenseffort = loweffort = mediumeffort = high

1 / 6 providers expose no reasoning control (reasoning_options: []).

Related models

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$0.70$02026-08-052026-08-132026-08-05 · Input list $0.182026-08-06 · Input list $0.182026-08-07 · Input list $0.182026-08-08 · Input list $0.182026-08-09 · Input list $0.182026-08-10 · Input list $0.182026-08-11 · Input list $0.182026-08-12 · Input list $0.182026-08-13 · Input list $0.0722026-08-05 · Output list $0.702026-08-06 · Output list $0.702026-08-07 · Output list $0.702026-08-08 · Output list $0.702026-08-09 · Output list $0.702026-08-10 · Output list $0.702026-08-11 · Output list $0.702026-08-12 · Output list $0.702026-08-13 · Output list $0.2872026-08-05 · Min blended $0.0612026-08-06 · Min blended $0.0612026-08-07 · Min blended $0.0612026-08-08 · Min blended $0.0612026-08-09 · Min blended $0.0612026-08-10 · Min blended $0.0612026-08-11 · Min blended $0.0612026-08-12 · Min blended $0.0612026-08-13 · Min blended $0.061