← Model list
Qwen3 4B
alibaba·alibaba/qwen3-4b-fp8·GA·Open weights
Qwen instruction model for multilingual chat, reasoning, and tool use
At a glance
Reference price$0.03 / $0.03 per 1M
Cache read — · NovitaAI
Lowest paid$0.03 / $0.03
NovitaAI Gateway
Context128,000
Output limit20,000
Capabilities
✓ ReasoningTool use? Structured output✓ TemperatureAttachments
Modalities
Text
Knowledge cutoff—
Released / updated2025-04-29 / 2025-04-29
Quality & performanceArtificial Analysis · Intelligence Index v4.1 · rep. tier Reasoning
Intelligence8.2
Value273
Coding—
Agentic—
Output speed—
TTFT—
Value formulaIQ 8.2 ÷ min blended $0.030 = 273
Reasoning tier → intelligence / speed (higher tier = stronger but slower)
Non-reasoningIQ 6.5
ReasoningIQ 8.2
Quality is independently evaluated by Artificial Analysis. Speed/latency are model-level medians.
Available at 1 providers1 with public prices
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| NovitaAI | Gateway | $0.03 | $0.03 | — | — | 128,000 | 20,000 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1NovitaAI$7.50
The cheapest paid channel is the only channel.
Benchmark
No upstream benchmark data for this model. For quality, see the Artificial Analysis intelligence score above.
Reasoning control
Upstream provides no control info
1 / 1 providers expose no reasoning control (reasoning_options: []).
Related models
GLM-4.7-Flashcheaper alternative$0 / $0Hy3cheaper alternative$0 / $0Nemotron 3 Nano 30B A3Bcheaper alternative$0 / $0Nemotron 3.5 Lightning 30B A3Bcheaper alternative$0 / $0
Price historyone sample accumulated per data sync
Input listOutput listMin blended