Skip to content
llm-spend
GitHub

DeepSeek

1M-token context, cheap direct pricing, and Microsoft Foundry resale markups from ~10% to a reported 4.5x — with one Data Zone lane that undercuts Global.

DeepSeek V4 Pro and V4 Flash ship real 1M-token windows (max output up to 384K). Pricing is a resale case study: the direct API is cheap, Microsoft Foundry resells it at a markup, and some Foundry tiers bill a cache meter the public page hides. Numbers below.

Both models are now tracked across all four Foundry lanes. V4 Pro and V4 Flash each have a first-party Global listing, a first-party Data Zone listing at the usual ~10% premium, and a Fireworks-hosted Data Zone listing. For V4 Pro the Fireworks lane is slightly the more expensive of the two Data Zone options; for V4 Flash it is dramatically the cheapest lane of all four, below even Global.

Pricing

per 1M tokens · USD / CHF
ModelTier / HostContextInputCachedOutputConfidence
DeepSeek-V4 Pro
1M context. Cache meter is published in Azure's retail catalog.
Foundry ·Global1M
384K max out
$1.74CHF 1.40$0.145CHF 0.117$3.48CHF 2.80official
Azure retail catalog: $1.74/M input, $0.145/M cached input, and $3.48/M output.
DeepSeek-V4 Flash
1M context. Cache meter is published in Azure's retail catalog.
Foundry ·Global1M
384K max out
$0.19CHF 0.153$0.028CHF 0.023$0.51CHF 0.411official
Azure retail catalog: $0.19/M input, $0.028/M cached input, and $0.51/M output.
DeepSeek-V4 Pro
First-party published cache-hit, cache-miss, and output rates after the 75% direct price cut. Peak/off-peak billing began 2026-08-16 16:00 UTC (peak 01:00-04:00 & 06:00-10:00 UTC, Monday to Friday only): off-peak $0.022/$0.66/$1.98, peak $0.044/$1.32/$3.96 — see the changelog for detail.
Direct API
DeepSeek direct API
1M$0.66CHF 0.531$0.022CHF 0.018$1.98CHF 1.59official
DeepSeek's own direct (non-cloud-resold) API pricing, including the published cached-input rate.
Off-peakChanges in 2h 20m
Peak$1.32 / $0.044 / $3.96Off-peak$0.66 / $0.022 / $1.98
DeepSeek-V4 Flash
First-party published cache-hit, cache-miss, and output rates. Peak/off-peak billing began 2026-08-16 16:00 UTC (peak 01:00-04:00 & 06:00-10:00 UTC, Monday to Friday only): off-peak $0.007/$0.22/$0.66, peak $0.014/$0.44/$1.32 — see the changelog for detail.
Direct API
DeepSeek direct API
1M$0.22CHF 0.177$0.007CHF 0.006$0.66CHF 0.531official
DeepSeek's own direct API pricing, including the published cached-input rate.
Off-peakChanges in 2h 20m
Peak$0.44 / $0.014 / $1.32Off-peak$0.22 / $0.007 / $0.66
DeepSeek-V4 Pro
Third-party inference provider's direct rate.
Direct API
Fireworks direct API
$1.74CHF 1.40$3.48CHF 2.80official
Fireworks direct API pricing.
DeepSeek-V4 Pro
First-party Foundry Data Zone deployment — slightly cheaper than the Fireworks-hosted Data Zone lane below ($1.925/$0.165/$3.828).
Foundry ·Data Zone1M$1.91CHF 1.54$0.16CHF 0.129$3.83CHF 3.08official
Azure Retail Prices API, product "Azure Deepseek Models": 'V4 Pro Inp DZ Tokens' $0.00191/1K, 'V4 Pro cached DZ Tokens' $0.00016/1K (effective 2026-07-01), 'V4 Pro Outp DZ Tokens' $0.00383/1K, consistent across 22 commercial regions. Captured 2026-07-27.
DeepSeek-V4 Pro
~10% above the Fireworks direct rate: the Data Zone premium. This is the Fireworks-hosted lane; a cheaper first-party Foundry Data Zone deployment also exists (see above).
Foundry ·Data Zone
Fireworks-hosted
$1.93CHF 1.55$0.165CHF 0.133$3.83CHF 3.08official
Fireworks official live pricing page.
DeepSeek-V4 Flash
First-party Foundry Data Zone deployment: the Global rate x1.10, rounded up to the meter's precision on every dimension. The Fireworks-hosted Data Zone lane below undercuts both this and Global.
Foundry ·Data Zone1M
384K max out
$0.21CHF 0.169$0.031CHF 0.025$0.56CHF 0.451official
Azure Retail Prices API, product "Azure Deepseek Models": 'V4 Flash Inp DZ Tokens' $0.00021/1K, 'V4 Flash cached DZ Tokens' $0.000031/1K, 'V4 Flash Outp DZ Tokens' $0.00056/1K, uniform across 22 commercial regions. Captured 2026-08-22. Each figure is the Global rate x1.10 rounded up ($0.19 to $0.209 to $0.21; $0.028 to $0.0308 to $0.031; $0.51 to $0.561 to $0.56). The two US-Gov regions price higher, as they do across this whole product; the commercial majority is used here. Microsoft separately publishes a 'V4 Flash 0731' meter set for the newer snapshot at $0.44/$0.014/$1.32 Global — this row tracks the plain 'V4 Flash' meters.
DeepSeek-V4 Flash
The cheapest DeepSeek lane on Foundry, and cheaper than the Global tier — an inversion of the usual pattern, because this is a different host undercutting the first-party listing rather than a tier discount.
Foundry ·Data Zone
Fireworks-hosted
$0.15CHF 0.121$0.03CHF 0.024$0.31CHF 0.25official
Azure Retail Prices API, product "Azure Fireworks Models": 'FW Deepseek-v4-Flash In DZ Tokens' $0.00015/1K, 'FW Deepseek-v4-Flash Cd In DZ Tokens' $0.00003/1K, 'FW Deepseek-v4-Flash Opt DZ Tokens' $0.00031/1K, effective 2026-08-01, uniform across all 20 commercial Data Zone regions with no US-Gov or rounding outlier. Captured 2026-08-22. Fireworks prices its own hosting, so this is not a multiple of DeepSeek's or Microsoft's rate, and the meter names no model snapshot.
Confidenceofficial published pagederived reconciled from billing estimate pattern-inferred
Field notes

Quirks & gotchas

Insight

Resale markup: ~10% to a reported 4.5x

Direct V4 Pro is $0.435 / CHF 0.35 input. A "Global" Microsoft Foundry V4 Pro has been reported at about 4.5x that, the widest markup here. The Fireworks Data Zone listing ($1.925 / CHF 1.55) is a milder ~10% over Fireworks direct ($1.74 / CHF 1.40).

Watch out

Azure cache meters are publicly priced

The Azure pricing summary does not show a cached-input column, but Azure's retail catalog lists one for both Global models: $0.145 / CHF 0.12 per M for V4 Pro (~91.7% off) and $0.028 / CHF 0.02 for V4 Flash (~85% off). Billing exports reconcile to those published meters.

Insight

V4 Flash: Data Zone can be cheaper than Global

Data Zone is normally the premium tier — it pins routing to a geography and charges about 10% for it. V4 Flash breaks that. Microsoft's own Data Zone deployment does follow the rule ($0.21 / CHF 0.17 input against $0.19 / CHF 0.15 Global), but the Fireworks-hosted Data Zone listing bills $0.15 / CHF 0.12 input and $0.31 / CHF 0.25 output — below the Global tier on every dimension, and the cheapest DeepSeek lane on Foundry.

The reason is that these are different sellers, not different tiers of one seller. Comparing tier labels across hosts tells you nothing about price; compare the meters. If you were going to accept Data Zone routing anyway, the Fireworks lane is strictly cheaper than staying on Global.

Note

1M context is a real autonomy lever

V4 Pro and V4 Flash carry 1M-token windows (max output up to 384K). Bigger windows mean less forced compaction, so less babysitting.

Note

No first-party embedding model

The chat models are generation-only. A separate deepseek-embedding-v2 (768-dim) exists, but for code RAG the Embeddings page recommends Cohere embed-v4.

Other providers
KimiGLMOpenAI / Azure OpenAIClaudeGeminiGrokQwenMistralMiniMaxEmbeddingsCompare all →