DeepSeek
1M-token context, cheap direct pricing, and Microsoft Foundry resale markups from ~10% to a reported 4.5x — with one Data Zone lane that undercuts Global.
DeepSeek V4 Pro and V4 Flash ship real 1M-token windows (max output up to 384K). Pricing is a resale case study: the direct API is cheap, Microsoft Foundry resells it at a markup, and some Foundry tiers bill a cache meter the public page hides. Numbers below.
Both models are now tracked across all four Foundry lanes. V4 Pro and V4 Flash each have a first-party Global listing, a first-party Data Zone listing at the usual ~10% premium, and a Fireworks-hosted Data Zone listing. For V4 Pro the Fireworks lane is slightly the more expensive of the two Data Zone options; for V4 Flash it is dramatically the cheapest lane of all four, below even Global.
Pricing
per 1M tokens · USD / CHF| Model | Tier / Host | Context | Input | Cached | Output | Confidence |
|---|---|---|---|---|---|---|
DeepSeek-V4 Pro 1M context. Cache meter is published in Azure's retail catalog. | Foundry ·Global | 1M 384K max out | $1.74CHF 1.40 | $0.145CHF 0.117 | $3.48CHF 2.80 | official Azure retail catalog: $1.74/M input, $0.145/M cached input, and $3.48/M output. |
DeepSeek-V4 Flash 1M context. Cache meter is published in Azure's retail catalog. | Foundry ·Global | 1M 384K max out | $0.19CHF 0.153 | $0.028CHF 0.023 | $0.51CHF 0.411 | official Azure retail catalog: $0.19/M input, $0.028/M cached input, and $0.51/M output. |
DeepSeek-V4 Pro First-party published cache-hit, cache-miss, and output rates after the 75% direct price cut. Peak/off-peak billing began 2026-08-16 16:00 UTC (peak 01:00-04:00 & 06:00-10:00 UTC, Monday to Friday only): off-peak $0.022/$0.66/$1.98, peak $0.044/$1.32/$3.96 — see the changelog for detail. | Direct API DeepSeek direct API | 1M | $0.66CHF 0.531 | $0.022CHF 0.018 | $1.98CHF 1.59 | official DeepSeek's own direct (non-cloud-resold) API pricing, including the published cached-input rate. |
Off-peakChanges in 2h 20m Peak$1.32 / $0.044 / $3.96Off-peak$0.66 / $0.022 / $1.98 | ||||||
DeepSeek-V4 Flash First-party published cache-hit, cache-miss, and output rates. Peak/off-peak billing began 2026-08-16 16:00 UTC (peak 01:00-04:00 & 06:00-10:00 UTC, Monday to Friday only): off-peak $0.007/$0.22/$0.66, peak $0.014/$0.44/$1.32 — see the changelog for detail. | Direct API DeepSeek direct API | 1M | $0.22CHF 0.177 | $0.007CHF 0.006 | $0.66CHF 0.531 | official DeepSeek's own direct API pricing, including the published cached-input rate. |
Off-peakChanges in 2h 20m Peak$0.44 / $0.014 / $1.32Off-peak$0.22 / $0.007 / $0.66 | ||||||
DeepSeek-V4 Pro Third-party inference provider's direct rate. | Direct API Fireworks direct API | — | $1.74CHF 1.40 | — | $3.48CHF 2.80 | official Fireworks direct API pricing. |
DeepSeek-V4 Pro First-party Foundry Data Zone deployment — slightly cheaper than the Fireworks-hosted Data Zone lane below ($1.925/$0.165/$3.828). | Foundry ·Data Zone | 1M | $1.91CHF 1.54 | $0.16CHF 0.129 | $3.83CHF 3.08 | official Azure Retail Prices API, product "Azure Deepseek Models": 'V4 Pro Inp DZ Tokens' $0.00191/1K, 'V4 Pro cached DZ Tokens' $0.00016/1K (effective 2026-07-01), 'V4 Pro Outp DZ Tokens' $0.00383/1K, consistent across 22 commercial regions. Captured 2026-07-27. |
DeepSeek-V4 Pro ~10% above the Fireworks direct rate: the Data Zone premium. This is the Fireworks-hosted lane; a cheaper first-party Foundry Data Zone deployment also exists (see above). | Foundry ·Data Zone Fireworks-hosted | — | $1.93CHF 1.55 | $0.165CHF 0.133 | $3.83CHF 3.08 | official Fireworks official live pricing page. |
DeepSeek-V4 Flash First-party Foundry Data Zone deployment: the Global rate x1.10, rounded up to the meter's precision on every dimension. The Fireworks-hosted Data Zone lane below undercuts both this and Global. | Foundry ·Data Zone | 1M 384K max out | $0.21CHF 0.169 | $0.031CHF 0.025 | $0.56CHF 0.451 | official Azure Retail Prices API, product "Azure Deepseek Models": 'V4 Flash Inp DZ Tokens' $0.00021/1K, 'V4 Flash cached DZ Tokens' $0.000031/1K, 'V4 Flash Outp DZ Tokens' $0.00056/1K, uniform across 22 commercial regions. Captured 2026-08-22. Each figure is the Global rate x1.10 rounded up ($0.19 to $0.209 to $0.21; $0.028 to $0.0308 to $0.031; $0.51 to $0.561 to $0.56). The two US-Gov regions price higher, as they do across this whole product; the commercial majority is used here. Microsoft separately publishes a 'V4 Flash 0731' meter set for the newer snapshot at $0.44/$0.014/$1.32 Global — this row tracks the plain 'V4 Flash' meters. |
DeepSeek-V4 Flash The cheapest DeepSeek lane on Foundry, and cheaper than the Global tier — an inversion of the usual pattern, because this is a different host undercutting the first-party listing rather than a tier discount. | Foundry ·Data Zone Fireworks-hosted | — | $0.15CHF 0.121 | $0.03CHF 0.024 | $0.31CHF 0.25 | official Azure Retail Prices API, product "Azure Fireworks Models": 'FW Deepseek-v4-Flash In DZ Tokens' $0.00015/1K, 'FW Deepseek-v4-Flash Cd In DZ Tokens' $0.00003/1K, 'FW Deepseek-v4-Flash Opt DZ Tokens' $0.00031/1K, effective 2026-08-01, uniform across all 20 commercial Data Zone regions with no US-Gov or rounding outlier. Captured 2026-08-22. Fireworks prices its own hosting, so this is not a multiple of DeepSeek's or Microsoft's rate, and the meter names no model snapshot. |
Quirks & gotchas
Resale markup: ~10% to a reported 4.5x
Direct V4 Pro is $0.435 / CHF 0.35 input. A "Global" Microsoft Foundry V4 Pro has been reported at about 4.5x that, the widest markup here. The Fireworks Data Zone listing ($1.925 / CHF 1.55) is a milder ~10% over Fireworks direct ($1.74 / CHF 1.40).
Azure cache meters are publicly priced
The Azure pricing summary does not show a cached-input column, but Azure's retail catalog lists one for both Global models: $0.145 / CHF 0.12 per M for V4 Pro (~91.7% off) and $0.028 / CHF 0.02 for V4 Flash (~85% off). Billing exports reconcile to those published meters.
V4 Flash: Data Zone can be cheaper than Global
Data Zone is normally the premium tier — it pins routing to a geography and charges about 10% for it. V4 Flash breaks that. Microsoft's own Data Zone deployment does follow the rule ($0.21 / CHF 0.17 input against $0.19 / CHF 0.15 Global), but the Fireworks-hosted Data Zone listing bills $0.15 / CHF 0.12 input and $0.31 / CHF 0.25 output — below the Global tier on every dimension, and the cheapest DeepSeek lane on Foundry.
The reason is that these are different sellers, not different tiers of one seller. Comparing tier labels across hosts tells you nothing about price; compare the meters. If you were going to accept Data Zone routing anyway, the Fireworks lane is strictly cheaper than staying on Global.
1M context is a real autonomy lever
V4 Pro and V4 Flash carry 1M-token windows (max output up to 384K). Bigger windows mean less forced compaction, so less babysitting.
No first-party embedding model
The chat models are generation-only. A separate deepseek-embedding-v2 (768-dim) exists, but for code RAG the Embeddings page recommends Cohere embed-v4.