LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

Qwen 3.5 397B

alibaba·alibaba/qwen3-5-397b-a17b·GA·Open weights·qwen series
Large open Qwen multimodal MoE for visual agents and long technical tasks

Specs & pricing

Input / output per 1M tokens
Reference price·Venice AI
$0.75 / $4.50
Blended $1.69 · Cache read —
Lowest paid·Ollama CloudCloud
$0.60 / $3.60
Blended $1.35 · 1.3× spread
Context
128,000
Output limit
32,768
Knowledge cutoff
—
Released / updated
2026-02-16 / 2026-06-11
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ Temperature✓ Attachments
Modalities
TextImageVideo

Available at 2 providers2 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
Ollama Cloud
qwen3.5:397b
Cloud$0.60$3.60——262,144 ⚠65,536
Venice AIGateway$0.75$4.50——128,00032,768

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

Toggle (on / off)effort = noneeffort = loweffort = mediumeffort = high

Interleaved thinking (reasoning between tool calls) is declared by 1 of 2 providers.

Your usage cost

1Ollama Cloud$300.00
2Venice AI$375.00
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input list $0.75Output list $4.50Min blended $1.35

Price history accumulates from each data sync; currently only 1 sample(s) (2026-09-24). Each future sync adds a point, and once accumulated a line is drawn here.

Related models

Qwen-Image-2.1same series$0 / $0Qwen 3.8 27B Hemingwaysame series$0.25 / $1.50Qwen 3.8 27B Cybersecuritysame series$0.10 / $0.60GLM-5.3-Flashcheaper alternative$0.15 / $0.50DeepSeek V4.1 Flashcheaper alternative$0.15 / $0.60DeepSeek V4 Flash 0731cheaper alternative$0.45 / $1.34GPT-5.6 Lunacheaper alternative$0.20 / $1.20
Data partly from models.dev (MIT) · Venice AI official docs ↗