OpenAI / Azure OpenAI
Azure normally resells OpenAI 1:1, but a promotional cut on the GPT-5.6 Sol flagship has opened a fresh gap. Deployment type and Responses-API-only variants are the other catches.
Azure OpenAI has historically matched OpenAI's direct pricing 1:1, so no resale markup. What changes is the deployment type on Microsoft Foundry (Global, Data Zone, Regional; see below). GPT-5.6 (Sol / Terra / Luna) hit GA on 2026-07-09 and has official Azure Foundry meters covering cached-input and cache-write plus Data Zone (+10%) and long-context tiers.
That 1:1 parity is not holding. OpenAI cut Terra and Luna on 2026-07-30 and Foundry took three weeks to follow, with a meter tranche effective 2026-08-01 that restored parity on those two. Then on 2026-08-21 OpenAI cut the Sol flagship — to $4.00/$0.40/$20.00 short context and $8.00/$0.80/$30.00 long context, described on its pricing page as promotional and available "at least through November 21, 2026" — and the Foundry Sol meters have not moved. Every Sol row below is the Azure meter, still on its original 2026-07-01 tranche, which now runs 1.25x OpenAI's direct input rate and 1.50x its direct output rate. Terra and Luna remain at parity.
Pricing
per 1M tokens · USD / CHF| Model | Tier / Host | Input | Cached | Output | Confidence |
|---|---|---|---|---|---|
GPT-5.6 Sol Flagship (hardest reasoning / coding / agentic). GA 2026-07-09. No longer at parity with OpenAI direct: OpenAI cut Sol to $4.00/$0.40/$20.00 on 2026-08-21 as promotional pricing, and this Foundry meter has not followed — Foundry is 1.25x direct on input and cached input, and 1.50x on output. | Foundry ·Global | $5.00CHF 4.03 | $0.50CHF 0.403 | $30.00CHF 24.15 | official Azure Retail Prices API 'Foundry Models' meters (5.6 sol Std Gl, effective 2026-07-01; captured 2026-07-21, re-verified unchanged in a full paged sweep on 2026-08-22 — all sol meters still carry the single 2026-07-01 tranche, with no 2026-08-01 tranche of the kind that carried the Terra and Luna cuts). Cache write bills at 1.25x uncached input ($6.25/M meter); reads stay ~90% off. OpenAI's own rate for gpt-5.6-sol, read via raw DOM from developers.openai.com/api/docs/pricing on 2026-08-22, is $4.00/$0.40/$20.00 with a $5.00/M cache write, labelled promotional and "available at least through November 21, 2026". |
Priority$10.00 / $1.00 / $60.00 | |||||
GPT-5.6 Sol ~10% Data Zone premium over Global. Because the Global Sol meter did not follow OpenAI's 2026-08-21 promotional cut, this row is roughly 1.38x OpenAI's direct input rate and 1.65x its direct output rate. | Foundry ·Data Zone | $5.50CHF 4.43 | $0.55CHF 0.443 | $33.00CHF 26.57 | official Azure Retail Prices API 'Foundry Models' meters (5.6 sol Std DZ, effective 2026-07-01; captured 2026-07-21, re-verified unchanged 2026-08-22 across the 13 US/EU Data Zone regions). |
GPT-5.6 Sol Long Context Long Context tier; all meters roughly double the short-context rates. Also left behind by OpenAI's 2026-08-21 promotional Sol cut, which took the direct long-context rate to $8.00/$0.80/$30.00 — so Foundry is 1.25x direct on input and 1.50x on output here too. | Foundry ·Global | $10.00CHF 8.05 | $1.00CHF 0.805 | $45.00CHF 36.23 | official Azure Retail Prices API 'Foundry Models' meters (5.6 sol LongCo Std Gl, effective 2026-07-01; captured 2026-07-21, re-verified unchanged 2026-08-22). OpenAI's direct long-context rate for gpt-5.6-sol, read via raw DOM on 2026-08-22, is $8.00/$0.80/$30.00 with a $10.00/M cache write. |
GPT-5.6 Terra Balanced production tier. GA 2026-07-09. Cut 20% from $2.50/$0.25/$15.00 — Foundry matched OpenAI's 2026-07-30 direct-API cut with a meter tranche effective 2026-08-01, restoring 1:1 parity. Cache write bills at $2.50/M (1.25x input). | Foundry ·Global | $2.00CHF 1.61 | $0.20CHF 0.161 | $12.00CHF 9.66 | official Azure Retail Prices API 'Foundry Models', full paged sweep captured 2026-08-20: '5.6 terra ShortCo Inp Std Gl 1M Tokens' $2.00/M, 'Cd Inp Std Gl' $0.20/M, 'Cd Wr Std Gl' $2.50/M, 'Opt Std Gl' $12.00/M — all effective 2026-08-01, and the superseded 2026-07-01 tranche ($2.50/$0.25/$15.00) no longer appears for any terra meter. Matches OpenAI's direct rate exactly (developers.openai.com/api/docs/pricing, read via raw DOM the same day). Data Zone meters moved with it, to $2.20/$0.22/$13.20 across 13 US/EU regions. |
Priority$4.00 / $0.40 / $24.00 | |||||
GPT-5.6 Terra Long Context Long Context tier, for prompts past the short-context threshold; input and cached input are 2x the short-context rates, output 1.5x. Cut from $5.00/$0.50/$22.50 in the same 2026-08-01 Foundry tranche as the short-context Terra row, matching OpenAI's direct long-context rate exactly. | Foundry ·Global | $4.00CHF 3.22 | $0.40CHF 0.322 | $18.00CHF 14.49 | official Azure Retail Prices API 'Foundry Models', full paged sweep captured 2026-08-20: '5.6 terra LongCo Inp Std Gl 1M Tokens' $4.00/M, 'Cd Inp Std Gl' $0.40/M, 'Cd Wr Std Gl' $5.00/M, 'Opt Std Gl' $18.00/M — effective 2026-08-01, with no 2026-07-01 terra tranche remaining. Equals OpenAI's published direct long-context rate for gpt-5.6-terra ($4.00/$0.40/$18.00), read via raw DOM from developers.openai.com/api/docs/pricing the same day. No LongCo priority (PP) meter exists for any GPT-5.6 variant. |
GPT-5.6 Luna Fast / cheap, high-volume. GA 2026-07-09. Cut 80% from $1.00/$0.10/$6.00 — Foundry matched OpenAI's 2026-07-30 direct-API cut with a meter tranche effective 2026-08-01, restoring 1:1 parity. Now the cheapest tracked lane on Foundry. Cache write bills at $0.25/M (1.25x input). | Foundry ·Global | $0.20CHF 0.161 | $0.02CHF 0.016 | $1.20CHF 0.966 | official Azure Retail Prices API 'Foundry Models', full paged sweep captured 2026-08-20: '5.6 luna ShortCo Inp Std Gl 1M Tokens' $0.20/M, 'Cd Inp Std Gl' $0.02/M, 'Cd Wr Std Gl' $0.25/M, 'Opt Std Gl' $1.20/M — all effective 2026-08-01, and the superseded 2026-07-01 tranche ($1.00/$0.10/$6.00) no longer appears for any luna meter. Matches OpenAI's direct rate exactly (developers.openai.com/api/docs/pricing, read via raw DOM the same day). Data Zone meters moved with it, to $0.22/$0.022/$1.32 across 13 US/EU regions. |
Priority$0.40 / $0.04 / $2.40 | |||||
GPT-5.6 Luna Long Context Long Context tier, for prompts past the short-context threshold; input and cached input are 2x the short-context rates, output 1.5x. Cut from $2.00/$0.20/$9.00 in the same 2026-08-01 Foundry tranche as the short-context Luna row, matching OpenAI's direct long-context rate exactly. | Foundry ·Global | $0.40CHF 0.322 | $0.04CHF 0.032 | $1.80CHF 1.45 | official Azure Retail Prices API 'Foundry Models', full paged sweep captured 2026-08-20: '5.6 luna LongCo Inp Std Gl 1M Tokens' $0.40/M, 'Cd Inp Std Gl' $0.04/M, 'Cd Wr Std Gl' $0.50/M, 'Opt Std Gl' $1.80/M — effective 2026-08-01, with no 2026-07-01 luna tranche remaining. Equals OpenAI's published direct long-context rate for gpt-5.6-luna ($0.40/$0.04/$1.80), read via raw DOM from developers.openai.com/api/docs/pricing the same day. No LongCo priority (PP) meter exists for any GPT-5.6 variant. |
GPT-5.5 | Foundry ·Global | $5.00CHF 4.03 | $0.50CHF 0.403 | $30.00CHF 24.15 | official Azure OpenAI pricing page. |
GPT-5.5 This is the US/EU data zone rate (1.10x Global). APAC data-zone regions (australiaeast, centralindia, eastasia, japaneast, japanwest, jioindiawest, koreacentral, southeastasia, southindia) price higher, at 1.20x Global: $6.00/M input, $0.60/M cached, $36.00/M output. | Foundry ·Data Zone | $5.50CHF 4.43 | $0.55CHF 0.443 | $33.00CHF 26.57 | official Azure OpenAI pricing page. |
GPT-5.5 Long Context Long Context tier. | Foundry ·Global | $10.00CHF 8.05 | $1.00CHF 0.805 | $45.00CHF 36.23 | official Azure OpenAI pricing page. |
GPT-5.3 Codex / Chat The -codex variant is Responses-API only. | Foundry ·Global | $1.75CHF 1.41 | $0.175CHF 0.141 | $14.00CHF 11.27 | official Azure OpenAI pricing page. Cached input re-verified against the Azure Retail Prices API (serviceName 'Foundry Models'), captured 2026-07-26: exact meter value is $0.175/M, correcting the earlier rounded $0.18/M. |
GPT-5.2 / Codex The -codex variant is Responses-API only. | Foundry ·Global | $1.75CHF 1.41 | $0.175CHF 0.141 | $14.00CHF 11.27 | official Azure OpenAI pricing page. Cached input re-verified against the Azure Retail Prices API (serviceName 'Foundry Models'), captured 2026-07-26: exact meter value is $0.175/M, correcting the earlier rounded $0.18/M. |
GPT-5.2 / Codex This is the US/EU data zone rate (1.10x Global). APAC data-zone regions (australiaeast, centralindia, eastasia, japaneast, japanwest, jioindiawest, koreacentral, southeastasia, southindia) price higher, at 1.20x Global: $2.10/M input, $0.21/M cached, $16.80/M output. | Foundry ·Data Zone | $1.93CHF 1.55 | $0.193CHF 0.155 | $15.40CHF 12.40 | official Azure OpenAI pricing page. Input and cached input re-verified against the Azure Retail Prices API (serviceName 'Foundry Models'), captured 2026-07-26: exact meter values are $1.925/M input and $0.1925/M cached input, correcting the earlier rounded $1.93/$0.2. Output ($15.40/M) was already exact. |
Quirks & gotchas
Deployment types: Global vs Data Zone vs Regional
Global routes to any datacenter: cheapest, highest throughput. Data Zone pins routing to US or EU and adds ~10%. Regional pins to one region and is the most restrictive and priciest. Pick Global unless data residency forces otherwise. The premium buys geography, not capability.
Data Zone is two prices, not one — APAC costs more
For Microsoft's first-party OpenAI lines, "Data Zone" isn't a single premium. Two disjoint region sets carry different rates: US/EU data zone (centralus, eastus, eastus2, francecentral, germanywestcentral, northcentralus, polandcentral, southcentralus, spaincentral, swedencentral, westeurope, westus, westus3, and on some models northeurope) bills at exactly 1.10x Global, while APAC data zone (australiaeast, centralindia, eastasia, japaneast, japanwest, jioindiawest, koreacentral, southeastasia, southindia) bills at exactly 1.20x Global. The APAC rows are effective 2026-06-01, added roughly six months after the US/EU rows for GPT-5.2.
Concretely: GPT-5.2 Data Zone runs $1.925/$0.1925/$15.40 in US/EU but $2.10/$0.21/$16.80 in APAC; GPT-5.5 short-context runs $5.50/$0.55/$33.00 in US/EU but $6.00/$0.60/$36.00 in APAC (long-context scales the same way). A customer deploying in an APAC region pays a real 20% premium over Global, not the 10% the catalog's single Data Zone row implies — budget accordingly.
GPT-5.3 chat and the entire GPT-5.6 family (Sol/Terra/Luna) have no APAC data-zone rows at all yet — Data Zone there is still a single US/EU-only price. And this split is exclusive to Microsoft's first-party OpenAI-hosted lines: Grok, Kimi, GLM, MiniMax, Mistral, and DeepSeek (both native and Fireworks-hosted) all publish a single Data Zone rate with no APAC surcharge. Captured from the Azure Retail Prices API on 2026-07-27.
A third split is now appearing, and it is the one to watch. Microsoft announced that from 2026-09-01 the EU data zone rises from 1.10x to 1.20x, applying only to models Foundry launches on or after that date — existing models are grandfathered, so nothing in this catalog changes on that day. The first meter with that shape has already shipped: a newly listed gpt-5-chat-latest snapshot bills its data-zone output at $33.00/M in seven US regions (1.10x Global) but $36.00/M in six EU regions (1.20x) — francecentral, germanywestcentral, polandcentral, spaincentral, swedencentral and westeurope. Expect new EU data-zone deployments to cost 20% over Global, not 10%. Captured from the Azure Retail Prices API on 2026-08-20.
-codex variants are Responses-API only
GPT-5.3-Codex, GPT-5.2-Codex and other "-codex" variants only support the Responses API (/v1/responses), not Chat Completions. Clients that default to Chat Completions fail with a 400 "unsupported operation" until reconfigured.
GPT-5.6 cache-write billing changed
GPT-5.6 bills cache writes at 1.25x the uncached input rate (was the standard input rate). Reads stay ~90% off. Factor the write premium into high-churn prompts.
Sol is now the model Foundry has not repriced
OpenAI cut GPT-5.6 Terra and Luna on 2026-07-30, and for three weeks the Azure Foundry meters did not follow. That gap closed on 2026-08-01, when a new Foundry tranche picked up the cut rates on Global, Data Zone, long-context and priority meters alike. Terra and Luna are at 1:1 with OpenAI direct today.
The same thing has now happened to the flagship. On 2026-08-21 OpenAI cut GPT-5.6 Sol to $4.00/$0.40/$20.00 short context and $8.00/$0.80/$30.00 long context — its pricing page calls this promotional and "available at least through November 21, 2026" — while every Sol meter in the Azure retail catalog still sits on its original 2026-07-01 tranche at $5.00/$0.50/$30.00. Running Sol on Foundry Global costs 25% more per input token and 50% more per output token than going direct to OpenAI, and Data Zone stacks its usual 10% on top of that.
Two things follow. If you are on Foundry for Sol specifically and have no data-residency requirement, the direct API is materially cheaper for as long as the promotion runs. And because OpenAI framed the cut as promotional with an open-ended "at least through" date rather than a fixed reversion, no reversion rate is published — this catalog does not stage a future price it cannot cite, so the Sol rows track the Azure meter and the gap is described here instead.
Worth remembering: when Terra and Luna lagged, Azure support answers and at least one downstream cost tracker described the cut as already applied on Azure while the retail catalog still billed the old rate. The retail meter is what bills you.
Benchmark leader, real-world laggard
GPT-5.3-Codex tops coding benchmarks but can lag in real agentic use, with poor context retention, versus DeepSeek V4 Pro and GLM-5.2. Likely a smaller window and less mature Responses-API support in some clients.