Skip to content
llm-spend
GitHub
Cost anatomy

DeepSeek-V4 Flash

DeepSeekFoundry ·Data Zone

Fresh input / cached input / output cost breakdown, active and scheduled rates, provenance, same-model deployment markup, and cost-comparable alternatives for this purchasable lane.

Start with a pattern

Illustrative workloads

Pick a useful starting shape, then fine-tune every input below.

Fine-tune

Workload and pricing scenario

Adjust the token shape, cache behavior, time window, and service tier. Every number below updates immediately.

Interactive

Workload cost calculator

60M
210K
90%

Share of input tokens served from cache. Applied only to models with a cache meter.

Time

Rows with a rate that changes by time of day, promo window or service tier are priced for this scenario instead of their flat listed rate; a brand-colored label under the model name names which one applies. Context size for band-priced rows uses the input-token count below.

Picking a specific hour also previews time-of-day rates that are scheduled but have not started billing yet; those rows are labelled · preview with their start instant. Now only ever shows rates that are billable at this moment.

Rate state

What this lane charges right now

Priced under the workload and scenario above. The full published schedule below always shows every variant regardless of which scenario is selected.

Input / 1M
$0.21
CHF 0.169
Cached input / 1M
$0.031
CHF 0.025
Output / 1M
$0.56
CHF 0.451
Cost anatomy

Where this workload's cost goes

60M input tokens (90% cache hit) and 210K output tokens, at the resolved rate above.

$1.26
Fresh input · CHF 1.01
$1.67
Cached input · CHF 1.35
$0.118
Output · CHF 0.095
$3.05
Workload total · CHF 2.46

$1.26 + $1.67 + $0.118 = $3.05

Provenance

How much to trust these numbers

InputofficialCached inputofficialOutputofficial

First-party Foundry Data Zone deployment: the Global rate x1.10, rounded up to the meter's precision on every dimension. The Fireworks-hosted Data Zone lane below undercuts both this and Global.

Azure Retail Prices API, product "Azure Deepseek Models": 'V4 Flash Inp DZ Tokens' $0.00021/1K, 'V4 Flash cached DZ Tokens' $0.000031/1K, 'V4 Flash Outp DZ Tokens' $0.00056/1K, uniform across 22 commercial regions. Captured 2026-08-22. Each figure is the Global rate x1.10 rounded up ($0.19 to $0.209 to $0.21; $0.028 to $0.0308 to $0.031; $0.51 to $0.561 to $0.56). The two US-Gov regions price higher, as they do across this whole product; the commercial majority is used here. Microsoft separately publishes a 'V4 Flash 0731' meter set for the newer snapshot at $0.44/$0.014/$1.32 Global — this row tracks the plain 'V4 Flash' meters.

Direct vs Foundry

Same-model deployment markup

How much this Foundry lane costs over the same model's Direct API rate.

LaneDeploymentWorkload costFoundry markup
DeepSeek-V4 Flash
DeepSeek direct API
Direct API$1.84+$1.22 (+66%)
Cost-comparable

Other lanes within ±25% of this workload's cost

Same provider first, then ranked by cost proximity, for 60M in / 210K out. This is a cost ranking only — never a quality recommendation.

  • DeepSeek-V4 Flash
    DeepSeekSame provider
    Foundry ·Global
    $2.76
    CHF 2.22
    -$0.292 (-10%) vs this lane
  • DeepSeek-V4 Flash
    DeepSeek · Fireworks-hostedSame provider
    Foundry ·Data Zone
    $2.59
    CHF 2.08
    -$0.467 (-15%) vs this lane
  • Qwen3.6 Flash
    Qwen · Model Studio (Intl)
    Direct API
    $3.17
    CHF 2.55
    +$0.113 (+4%) vs this lane
  • GPT-5.6 Luna
    OpenAI / Azure OpenAI
    Foundry ·Global
    $2.53
    CHF 2.04
    -$0.52 (-17%) vs this lane