← Model list

Qwen3.8 Max Preview

alibaba·alibaba/qwen3.8-max-preview·BETA·Closed·qwen series·NEW

Preview Qwen flagship for million-token multimodal reasoning and long-horizon agentic workflows

At a glance
Reference price$2.50 / $7.50 per 1M
Cache read · Impossibl
Lowest paid$1.81 / $5.45
DevPass (LLM Gateway) Gateway · 1.4× spread
2 more $0 channels
Context1,000,000
Output limit131,072
Capabilities
ReasoningTool useStructured outputTemperatureAttachments
Modalities
TextImageVideo
Knowledge cutoff
Released / updated2026-07-19 / 2026-07-19

Quality & performance

Artificial Analysis doesn't cover this model (267 of 2059 have data). Quality data comes from independent evals covering widely used models.

Available at 5 providers3 with public prices · 2 subscription-covered

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
Alibaba Token PlanGatewaySubscription$0$01,000,000131,072beta
Alibaba Token Plan (China)GatewaySubscription$0$01,000,000131,072beta
DevPass (LLM Gateway)
qwen3.8-max
Gateway$1.81$5.45$0.21$2.501,000,0001,000,000
Charm Hyper
qwen3.8-max
Gateway$2.00$6.00$0.251,000,00065,536
ImpossiblGateway$2.50$7.501,000,000131,072

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Your usage cost

1DevPass (LLM Gateway)$442.71
2Charm Hyper$490.00
3Impossibl$875.00

2 more channels offer $0 (Alibaba Token Plan, Alibaba Token Plan (China)); free tiers usually have rate limits and no SLA, excluded from ranking.

The cheapest paid channel is the only channel.

Benchmark15 items

NameConditionsScoreMetricSource
Terminal-Benchvariant: xhigh · v2.186.6accuracySource ↗
SWE-Bench Proharness: Claude Code · variant: xhigh67.7resolve rateSource ↗
DeepSWEharness: Claude Code · variant: xhigh · v1.156.6resolve rateSource ↗
NL2Repoharness: Claude Code · variant: xhigh55.9resolve rateSource ↗
FrontierSWEharness: Claude Code · variant: xhigh73.5dominance scoreSource ↗
MLS-Bench-Liteharness: Claude Code · variant: xhigh41scoreSource ↗
AutomationBenchvariant: xhigh · dataset: 600-task public subset27.3pass@1Source ↗
Toolathlon Verifiedvariant: xhigh72.5pass@1Source ↗
WideSearchvariant: xhigh81.9F1Source ↗
Humanity's Last Examvariant: xhigh, with tools56.2accuracySource ↗
GPQA Diamondvariant: xhigh92.6accuracySource ↗
Humanity's Last Examvariant: xhigh, no tools43.6accuracySource ↗
IFBenchvariant: xhigh82.8scoreSource ↗
OSWorld-Verifiedvariant: xhigh86.1success rateSource ↗
MMMU Provariant: xhigh82.3accuracySource ↗

The same benchmark scores very differently across harness / dataset, so the qualifying conditions must be shown together.

Reasoning control

effort = loweffort = mediumeffort = xhighbudget_tokens

1 / 5 providers expose no reasoning control (reasoning_options: []).

Related models

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$7.50$02026-08-052026-08-132026-08-05 · Input list $2.502026-08-06 · Input list $2.502026-08-07 · Input list $2.502026-08-08 · Input list $2.502026-08-09 · Input list $2.502026-08-10 · Input list $2.502026-08-11 · Input list $2.502026-08-12 · Input list $2.502026-08-13 · Input list $2.502026-08-05 · Output list $7.502026-08-06 · Output list $7.502026-08-07 · Output list $7.502026-08-08 · Output list $7.502026-08-09 · Output list $7.502026-08-10 · Output list $7.502026-08-11 · Output list $7.502026-08-12 · Output list $7.502026-08-13 · Output list $7.502026-08-05 · Min blended $2.382026-08-06 · Min blended $2.382026-08-07 · Min blended $2.382026-08-08 · Min blended $2.382026-08-09 · Min blended $2.382026-08-10 · Min blended $2.382026-08-11 · Min blended $2.382026-08-12 · Min blended $2.382026-08-13 · Min blended $2.72