Compare cost, not sticker price
Shape a workload and see the lowest-cost purchasable lanes across Microsoft Foundry and direct APIs. Every result stays traceable to the complete rate table below.
Illustrative workloads
Pick a useful starting shape, then fine-tune every input below.
Lowest cost for this workload
Workload and pricing scenario
Adjust the token shape, cache behavior, time window, and service tier. Every signal updates immediately.
Workload cost calculator
Share of input tokens served from cache. Applied only to models with a cache meter.
Rows with a rate that changes by time of day, promo window or service tier are priced for this scenario instead of their flat listed rate; a brand-colored label under the model name names which one applies. Context size for band-priced rows uses the input-token count below.
Picking a specific hour also previews time-of-day rates that are scheduled but have not started billing yet; those rows are labelled · preview with their start instant. Now only ever shows rates that are billable at this moment.
Focus the catalog
| Shortlist | ||||||||
|---|---|---|---|---|---|---|---|---|
| Qwen | Qwen3.7 Flash Model Studio (Intl) | Direct API | $0.10 | $0.01† | $0.40 | $0.019 | $1.22 CHF 0.985 | |
| GLM | GLM-5.3-Flash Z.ai direct API | Direct API | $0.075 | $0.015 | $0.25 | $0.021 | $1.31 CHF 1.06 | |
| DeepSeek | DeepSeek-V4 Flash DeepSeek direct API Off-peak | Direct API | $0.22 | $0.007 | $0.66 | $0.028 | $1.84 CHF 1.48 | |
| Qwen | Qwen3.8 Flash Model Studio (Intl) | Direct API | $0.15 | $0.016 | $0.47 | $0.029 | $1.86 CHF 1.50 | |
| OpenAI / Azure OpenAI | GPT-5.6 Luna | Foundry ·Global | $0.20 | $0.02 | $1.20 | $0.038 | $2.53 CHF 2.04 | |
| DeepSeek | DeepSeek-V4 Flash Fireworks-hosted | Foundry ·Data Zone | $0.15 | $0.03 | $0.31 | $0.042 | $2.59 CHF 2.08 | |
| DeepSeek | DeepSeek-V4 Flash | Foundry ·Global | $0.19 | $0.028 | $0.51 | $0.044 | $2.76 CHF 2.22 | |
| DeepSeek | DeepSeek-V4 Flash | Foundry ·Data Zone | $0.21 | $0.031 | $0.56 | $0.049 | $3.05 CHF 2.46 | |
| Qwen | Qwen3.6 Flash Model Studio (Intl) | Direct API | $0.25 | $0.025† | $1.50 | $0.048 | $3.17 CHF 2.55 | |
| Qwen | Qwen3.7 Plus (Promo) Model Studio (Intl) | Direct API | $0.32 | $0.032† | $1.28 | $0.061 | $3.92 CHF 3.15 | |
| MiniMax | MiniMax M2.5 Fireworks-hosted | Foundry ·Data Zone | $0.33 | $0.033 | $1.32 | $0.063 | $4.04 CHF 3.25 | |
| OpenAI / Azure OpenAI | GPT-5.6 Luna Long Context | Foundry ·Global | $0.40 | $0.04 | $1.80 | $0.076 | $4.94 CHF 3.98 |
Tier badges reading Foundry · … are Microsoft Foundry deployment tiers (Global routes to any datacenter, Data Zone pins to US or EU at roughly a 10% premium, Regional pins to one region); Direct API is the model developer’s own first-party API. Blended in is the effective $/1M input paid after the cache split. * marks models with no cache meter (hit rate ignored). Daggers † /‡ mark derived / estimated rates. A brand-colored label under a model name names the variant priced under the selected scenario, when it differs from the row’s base rate. A · preview label means the opposite: that rate is scheduled but is not billing yet, and is only shown because a specific hour was picked above — the row still charges its base rate until the stated start instant. The default Now scenario never shows a preview. The ☆ control pins a lane to the shortlist tray, which ranks pinned lanes by cost for the selected workload only.