LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

Qwen3.8 2.4T A95B (NVFP4)

alibaba·alibaba/qwen3.8-2.4t-a95b-nvfp4·GA·Open weights·qwen series·NEW
Open-weight sparse MoE (2.4T total, 95B active), the open-weight twin of Qwen3.8 Max for coding, research, complex reasoning, and agentic workflows

Specs & pricing

Input / output per 1M tokens
Reference price·RunInfra
$2.00 / $6.00
Blended $3.00 · Cache read $0.20
Lowest paid·RunInfraGateway
$2.00 / $6.00
Blended $3.00
Context
262,144
Output limit
32,768
Knowledge cutoff
—
Released / updated
2026-08-12 / 2026-08-12
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
Modalities
Text

Available at 1 providers1 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
RunInfra
Inferact/Qwen3.8-2.4T-A95B-NVFP4
Gateway$2.00$6.00$0.20—262,14432,768

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

effort = loweffort = mediumeffort = xhigh

Your usage cost

1RunInfra$484.00
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input list $2.00Output list $6.00Min blended $3.00

Price history accumulates from each data sync; currently only 1 sample(s) (2026-08-17). Each future sync adds a point, and once accumulated a line is drawn here.

Related models

Qwen-Image-2.1same series$0 / $0Qwen 3.8 27B Hemingwaysame series$0.25 / $1.50Qwen 3.8 27B Cybersecuritysame series$0.10 / $0.60GLM-5.3-Flashcheaper alternative$0.15 / $0.50DeepSeek V4.1 Flashcheaper alternative$0.15 / $0.60DeepSeek V4 Flash 0731cheaper alternative$0.45 / $1.34GPT-5.6 Lunacheaper alternative$0.20 / $1.20
Data partly from models.dev (MIT) · RunInfra official docs ↗