Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows
Qwen3.8 Max Preview by alibaba is offered by 6 providers on this page. Public prices are shown for 4 of them. Its reference price is $2.50 per 1M input tokens and $7.50 per 1M output tokens. The lowest paid channel is AIHubMix at $0.338 / $1.01 per 1M, about 7.4× below the reference price. That channel is a third-party gateway, so confirm its availability and rate limits before depending on it. It also has 2 covered by a paid subscription; free tiers usually carry rate limits, and subscription-covered access bills $0 per token only after the subscription fee.
The context window is 1,000,000 tokens, with an output limit of 131,072 tokens. It supports reasoning, tool use, and structured output. Accepted input modalities are Text, Image, and Video.
| Provider | Tier | Input | Output | Cache read | Cache write | Context | Output limit | Status |
|---|---|---|---|---|---|---|---|---|
| Alibaba Token Plan | Gateway | Subscription | $0 | $0 | 1,000,000 | 131,072 | deprecated | |
| Alibaba Token Plan (China) | Gateway | Subscription | $0 | $0 | 1,000,000 | 131,072 | deprecated | |
| AIHubMix | Gateway | $0.34 | $1.01 | $0.068 | $0.42 | 1,000,000 | 131,072 | |
| Charm Hyper qwen3.8-max | Gateway | $2.00 | $6.00 | $0.25 | — | 1,000,000 | 65,536 | |
| DevPass (LLM Gateway) qwen3.8-max | Gateway | $2.00 | $6.00 | $0.25 | $2.50 | 1,000,000 | 1,000,000 | |
| Impossibl | Gateway | $2.50 | $7.50 | — | — | 1,000,000 | 131,072 |
Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.
Interleaved thinking (reasoning between tool calls) is declared by 3 of 6 providers.
2 more channels offer $0 (Alibaba Token Plan, Alibaba Token Plan (China)); free tiers usually have rate limits and no SLA, excluded from ranking.
| Name | Conditions | Score | Metric | Source |
|---|---|---|---|---|
| Terminal-Bench | variant: xhigh · v2.1 | 86.6 | accuracy | Source ↗ |
| SWE-Bench Pro | harness: Claude Code · variant: xhigh | 67.7 | resolve rate | Source ↗ |
| DeepSWE | harness: Claude Code · variant: xhigh · v1.1 | 56.6 | resolve rate | Source ↗ |
| NL2Repo | harness: Claude Code · variant: xhigh | 55.9 | resolve rate | Source ↗ |
| FrontierSWE | harness: Claude Code · variant: xhigh | 73.5 | dominance score | Source ↗ |
| MLS-Bench-Lite | harness: Claude Code · variant: xhigh | 41 | score | Source ↗ |
| AutomationBench | variant: xhigh · dataset: 600-task public subset | 27.3 | pass@1 | Source ↗ |
| Toolathlon Verified | variant: xhigh | 72.5 | pass@1 | Source ↗ |
| WideSearch | variant: xhigh | 81.9 | F1 | Source ↗ |
| Humanity's Last Exam | variant: xhigh, with tools | 56.2 | accuracy | Source ↗ |
| GPQA Diamond | variant: xhigh | 92.6 | accuracy | Source ↗ |
| Humanity's Last Exam | variant: xhigh, no tools | 43.6 | accuracy | Source ↗ |
| IFBench | variant: xhigh | 82.8 | score | Source ↗ |
| OSWorld-Verified | variant: xhigh | 86.1 | success rate | Source ↗ |
| MMMU Pro | variant: xhigh | 82.3 | accuracy | Source ↗ |
The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.