LLM Pricing
PricingLeaderboardsToolsProvidersReleasesGuides

© 2026 LLM Pricing

About
·Contact
·Privacy
·RSS
← Model list

DeepSeek V4 Flash 0731 Fast

deepseek·deepseek/deepseek-v4-flash-0731-fast·GA·Open weights·deepseek-flash series·NEW
Fast DeepSeek model for efficient chat, coding help, and agent loops

Specs & pricing

Input / output per 1M tokens
Reference price·AIHubMix
$0.28 / $1.40
Blended $0.56 · Cache read $0.07
Lowest paid·Venice AIGateway
$0.35 / $0.70
Blended $0.44 · 1.2× spread
Context
1,000,000
Output limit
384,000
Knowledge cutoff
2025-05
Released / updated
2026-08-09 / 2026-08-11
Capabilities
✓ Reasoning✓ Tool use✓ Structured output✓ TemperatureAttachments
Modalities
Text

Available at 2 providers2 with public prices

ProviderTierInputOutputCache readCache writeContextOutput limitStatus
Venice AIGateway$0.35$0.70$0.088—1,000,00032,768
AIHubMixGateway$0.28$1.40$0.07—1,000,000384,000

Sorted by blended price (input×0.75 + output×0.25) asc. The official channel always shows regardless of rank. Whether a gateway's low price is actually usable can't be verified.

Reasoning control

effort = noneeffort = loweffort = higheffort = maxToggle (on / off)

Interleaved thinking (reasoning between tool calls) is declared by 1 of 2 providers.

Your usage cost

1Venice AI$73.50
2AIHubMix$100.80
The cheapest paid channel is the only channel.

Price historyone sample accumulated per data sync

Input listOutput listMin blended
$1.40$02026-08-122026-09-222026-08-12 · Input list $0.352026-08-13 · Input list $0.352026-09-22 · Input list $0.282026-08-12 · Output list $0.702026-08-13 · Output list $0.702026-09-22 · Output list $1.402026-08-12 · Min blended $0.442026-08-13 · Min blended $0.442026-09-22 · Min blended $0.44

Related models

DeepSeek Flash Latestsame series$0.003 / $2.40DeepSeek V4.1 Flashsame series$0.15 / $0.60deepseek-ai/DeepSeek-V4.1-Flash-Fastsame series$0.60 / $2.40GLM-5.3-Flashcheaper alternative$0.15 / $0.50Qwen3.8 Flashcheaper alternative$0.15 / $0.47GPT-6 Lunacheaper alternative$0.10 / $0.50
Data partly from models.dev (MIT) · AIHubMix official docs ↗