← Model list
Qwen3.8 Max Preview
alibaba·alibaba/qwen3.8-max-preview·BETA·Closed·qwen series·NEW
Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows
At a glance
Reference price$2.50 / $7.50 per 1M
Cache read — · Impossibl
Lowest paid$1.81 / $5.45
DevPass (LLM Gateway) Gateway · 1.4× spread
2 more $0 channels
Context1,000,000
Output limit131,072
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ Temperature✓ Attachments
Modalities
TextImageVideo
Knowledge cutoff—
Released / updated2026-07-19 / 2026-07-19
Quality & performance
Artificial Analysis doesn't cover this model (267 of 2059 have data). Quality data comes from independent evals covering widely used models.
Available at 5 providers3 with public prices · 2 subscription-covered
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Alibaba Token Plan | Gateway | Subscription | $0 | $0 | 1,000,000 | 131,072 | beta | |
| Alibaba Token Plan (China) | Gateway | Subscription | $0 | $0 | 1,000,000 | 131,072 | beta | |
| DevPass (LLM Gateway) qwen3.8-max | Gateway | $1.81 | $5.45 | $0.21 | $2.50 | 1,000,000 | 1,000,000 | |
| Charm Hyper qwen3.8-max | Gateway | $2.00 | $6.00 | $0.25 | — | 1,000,000 | 65,536 | |
| Impossibl | Gateway | $2.50 | $7.50 | — | — | 1,000,000 | 131,072 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Your usage cost
1DevPass (LLM Gateway)$442.71
2Charm Hyper$490.00
3Impossibl$875.00
2 more channels offer $0 (Alibaba Token Plan, Alibaba Token Plan (China)); free tiers usually have rate limits and no SLA, excluded from ranking.
The cheapest paid channel is the only channel.
Benchmark15 items
| Name | Conditions | Score | Metric | Source |
|---|---|---|---|---|
| Terminal-Bench | variant: xhigh · v2.1 | 86.6 | accuracy | Source ↗ |
| SWE-Bench Pro | harness: Claude Code · variant: xhigh | 67.7 | resolve rate | Source ↗ |
| DeepSWE | harness: Claude Code · variant: xhigh · v1.1 | 56.6 | resolve rate | Source ↗ |
| NL2Repo | harness: Claude Code · variant: xhigh | 55.9 | resolve rate | Source ↗ |
| FrontierSWE | harness: Claude Code · variant: xhigh | 73.5 | dominance score | Source ↗ |
| MLS-Bench-Lite | harness: Claude Code · variant: xhigh | 41 | score | Source ↗ |
| AutomationBench | variant: xhigh · dataset: 600-task public subset | 27.3 | pass@1 | Source ↗ |
| Toolathlon Verified | variant: xhigh | 72.5 | pass@1 | Source ↗ |
| WideSearch | variant: xhigh | 81.9 | F1 | Source ↗ |
| Humanity's Last Exam | variant: xhigh, with tools | 56.2 | accuracy | Source ↗ |
| GPQA Diamond | variant: xhigh | 92.6 | accuracy | Source ↗ |
| Humanity's Last Exam | variant: xhigh, no tools | 43.6 | accuracy | Source ↗ |
| IFBench | variant: xhigh | 82.8 | score | Source ↗ |
| OSWorld-Verified | variant: xhigh | 86.1 | success rate | Source ↗ |
| MMMU Pro | variant: xhigh | 82.3 | accuracy | Source ↗ |
The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.
Reasoning control
effort = loweffort = mediumeffort = xhighbudget_tokens
1 / 5 providers expose no reasoning control (reasoning_options: []).
Related models
Qwen3.5 0.8Bsame series$0.06 / $0.12Qwen3.5 4Bsame series$0.10 / $0.20Qwen3.8 27B TEEsame series$0.40 / $3.00GLM-5.2cheaper alternative$1.40 / $4.40DeepSeek V4 Procheaper alternative$0.435 / $0.87DeepSeek V4 Flashcheaper alternative$0.14 / $0.28DeepSeek V4 Flash 0731cheaper alternative$0.479 / $1.44
Price historyone sample accumulated per data sync
Input listOutput listMin blended