GLM
Zhipu / Z.aiGLM-5.3-Flash adds a multimodal, half-price direct lane, while GLM-5.3, 5.1 and 5.2 remain aligned on the full-rate tier.
GLM-5.3-Flash is now the newest and cheapest 5.x lane: Z.ai's native multimodal model is currently 50% off at $0.075/M input, $0.015/M cached input and $0.25/M output. The promotion ends at 24:00 on 2026-09-09 UTC+8, after which it returns to $0.15/$0.03/$0.50. It keeps a 1M-token window, 128K max output and always-enabled reasoning.
GLM-5.3 remains direct-API only so far with no Foundry meter published. It prices identically to 5.1 and 5.2 on all three dimensions ($1.40/M input, $0.26/M cached input, $4.40/M output), keeps the same 1M-token window with 128K max output, and always runs with reasoning enabled — three effort levels (low, high, max) rather than an on/off toggle.
GLM-5.2 lifts the window to a real 1M tokens (up 5x from 5.1's 200K), with 131K max output, the practical win for agentic coding. Input and output match 5.1's Data Zone rate, but Azure now publishes a dedicated 5.2 cached-input meter at $0.15/M, well below 5.1's $0.286/M (see below).
The original GLM-5 is still generally available and is the cheapest lane in this family: $1.10/M input and $3.52/M output on Foundry Data Zone, roughly 29% and 27% under 5.1 and 5.2, with the same 200K window as 5.1. If you do not need 5.2's 1M context, it is the value pick rather than a superseded model.
GLM-5-Turbo sits between GLM-5 and 5.1/5.2 at $1.20/M input and $4.00/M output, and is Z.ai-direct only with no Foundry meter.
Pricing
per 1M tokens · USD / CHF| Model | Tier / Host | Context | Input | Cached | Output | Confidence |
|---|---|---|---|---|---|---|
GLM 5 The cheapest GLM lane on Foundry: ~29% below 5.1/5.2 on input and ~27% below on output, for the same 200K window as 5.1. | Foundry ·Data Zone Fireworks-hosted | 200K 128K max out | $1.10CHF 0.886 | $0.22CHF 0.177 | $3.52CHF 2.83 | official Azure Retail Prices API 'FW GLM 5' meters, captured 2026-07-25 (effective 2026-06-01): input $0.0011/1K, output $0.00352/1K, cached input $0.00022/1K. Each is exactly 1.1x the Z.ai direct rate, the standard Data Zone premium. No Global-tier meter is published for GLM 5, matching 5.1 and 5.2. |
GLM 5.1 | Foundry ·Data Zone Fireworks-hosted | 200K | $1.54CHF 1.24 | $0.286CHF 0.23 | $4.84CHF 3.90 | official Fireworks official live pricing page. |
GLM 5.2 Cached input is officially published at $0.15/M, ~48% below GLM 5.1's $0.286/M. | Foundry ·Data Zone Fireworks-hosted | 1M 131K max out | $1.54CHF 1.24 | $0.15CHF 0.121 | $4.84CHF 3.90 | official Azure Retail Prices API 'FW GLM 5.2' meters, captured 2026-07-22 (effective 2026-07-01): input $0.00154/1K, output $0.00484/1K, cached input $0.00015/1K, uniform across regions. Earlier estimate (equal to GLM 5.1) was right on input/output but high on cache. |
GLM-5 Z.ai publishes a cached-input rate for GLM-5 ($0.20/M). Cached-input storage is currently free for a limited time, a billing dimension this catalog does not model. | Direct API Z.ai direct API | 200K 128K max out | $1.00CHF 0.805 | $0.20CHF 0.161 | $3.20CHF 2.58 | official Z.ai official pricing page, captured 2026-07-25: input $1, cached input $0.2, output $3.2 per 1M tokens. Model page lists a 200K window and 128K max output. The 'Limited-time Free' marker on that page applies only to the separate Cached Input Storage column, not to these rates. |
GLM-5-Turbo Agentic / tool-calling variant priced between GLM-5 and GLM-5.1/5.2; no Foundry meter, so Z.ai direct is the only lane. | Direct API Z.ai direct API | 200K 128K max out | $1.20CHF 0.966 | $0.24CHF 0.193 | $4.00CHF 3.22 | official Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-07-29: $1.20/M input, $0.24/M cached input, $4.00/M output. The model page (docs.z.ai/guides/llm/glm-5-turbo) lists a 200K window and 128K max output and positions it as optimized for agentic long-chain execution and high-throughput tool calling. A full sweep of Azure's Foundry catalog on 2026-07-29 found no GLM-5-Turbo meter — only GLM 5, 5.1 and 5.2 are resold there. |
GLM-5.1 Z.ai is the family's first-party API; input/output already matched Z.ai exactly, and the cache rate is now sourced from there too. | Direct API Z.ai direct API | 200K | $1.40CHF 1.13 | $0.26CHF 0.209 | $4.40CHF 3.54 | official Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-07-26: $1.40/M input, $0.26/M cached input, $4.40/M output. Vendor label corrected from 'Fireworks direct API' to 'Z.ai direct API' — Z.ai is the developer's own API per this catalog's Direct-tier definition, and previously had no cached-input figure recorded. |
GLM-5.2 Identical direct rate to GLM-5.1. Z.ai is the family's first-party API. | Direct API Z.ai direct API | 1M 131K max out | $1.40CHF 1.13 | $0.26CHF 0.209 | $4.40CHF 3.54 | official Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-07-26: $1.40/M input, $0.26/M cached input, $4.40/M output (5.1 and 5.2 priced identically). Vendor label corrected from 'Fireworks direct API' to 'Z.ai direct API', matching the GLM-5 Direct row and closing the family's mixed-vendor Direct lane. |
GLM-5.3 Prices identically to GLM-5.1 and GLM-5.2 on all three dimensions. Keeps the family's 1M-token window with 128K max output, and always runs with reasoning enabled — three effort levels (low, high, max); disabling reasoning is not supported. | Direct API Z.ai direct API | 1M 128K max out | $1.40CHF 1.13 | $0.26CHF 0.209 | $4.40CHF 3.54 | official Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-08-19: $1.40/M input, $0.26/M cached input, $4.40/M output per 1M tokens, USD — identical to GLM-5.1 and GLM-5.2. The model guide (docs.z.ai/guides/llm/glm-5.3) lists a 1M-token window, 128K max output, and reasoning always enabled with three effort levels (low, high, max). A full Azure Retail Prices API sweep on 2026-08-19 (29,405 rows) found no Foundry meter for GLM 5.3, so this direct-API row is the only lane. |
GLM-5.3-Flash Native multimodal GLM-5 model with image, video and file input. Reasoning is always enabled; the current 50% promotion covers input, cached input and output, then reverts to the published list rate on 2026-09-09 at 16:00 UTC. | Direct API Z.ai direct API | 1M 128K max out | $0.075CHF 0.06 | $0.015CHF 0.012 | $0.25CHF 0.201 | official Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-08-28: GLM-5.3-Flash is listed at $0.075/M input, $0.015/M cached input and $0.25/M output, with strikethrough list prices of $0.15/$0.03/$0.50. The page states that the 50% promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time). The official model guide (docs.z.ai/guides/vlm/glm-5.3-flash) lists image/video/text/file input, a 1M-token context window, 128K max output and reasoning that cannot be disabled. A full Azure Retail Prices API sweep on 2026-08-28 found no Foundry meter for GLM-5.3-Flash, so this direct-API row is the only lane. |
List price (from September 10) begins 9 Sep, 16:00 UTC List price (from September 10)$0.15 / $0.03 / $0.50 | ||||||
Quirks & gotchas
GLM 5's Data Zone rate is a clean 1.1x
All three of GLM 5's Foundry Data Zone meters land at exactly 1.1x the Z.ai direct rate: $1.00 to $1.10 input, $0.20 to $0.22 cached, $3.20 to $3.52 output. Two independently published sources agreeing on all three dimensions is the strongest confirmation this catalog gets that both numbers are right.
5.2 Data Zone: an estimate that held up
The unlisted GLM-5.2 Data Zone rate was set equal to GLM-5.1's ($1.54 / CHF 1.24 input, $0.286 / CHF 0.23 cached, $4.84 / CHF 3.90 output). Real invoices later matched almost exactly. A well-reasoned estimate, flagged as such, can hold until an official number lands.
1M context window on 5.2
GLM-5.2's 1M-token window (up from 200K on 5.1) puts it in the same autonomy tier as DeepSeek V4 for long runs.