Skip to content
llm-spend
GitHub

GLM

Zhipu / Z.ai

GLM-5.3-Flash adds a multimodal, half-price direct lane, while GLM-5.3, 5.1 and 5.2 remain aligned on the full-rate tier.

GLM-5.3-Flash is now the newest and cheapest 5.x lane: Z.ai's native multimodal model is currently 50% off at $0.075/M input, $0.015/M cached input and $0.25/M output. The promotion ends at 24:00 on 2026-09-09 UTC+8, after which it returns to $0.15/$0.03/$0.50. It keeps a 1M-token window, 128K max output and always-enabled reasoning.

GLM-5.3 remains direct-API only so far with no Foundry meter published. It prices identically to 5.1 and 5.2 on all three dimensions ($1.40/M input, $0.26/M cached input, $4.40/M output), keeps the same 1M-token window with 128K max output, and always runs with reasoning enabled — three effort levels (low, high, max) rather than an on/off toggle.

GLM-5.2 lifts the window to a real 1M tokens (up 5x from 5.1's 200K), with 131K max output, the practical win for agentic coding. Input and output match 5.1's Data Zone rate, but Azure now publishes a dedicated 5.2 cached-input meter at $0.15/M, well below 5.1's $0.286/M (see below).

The original GLM-5 is still generally available and is the cheapest lane in this family: $1.10/M input and $3.52/M output on Foundry Data Zone, roughly 29% and 27% under 5.1 and 5.2, with the same 200K window as 5.1. If you do not need 5.2's 1M context, it is the value pick rather than a superseded model.

GLM-5-Turbo sits between GLM-5 and 5.1/5.2 at $1.20/M input and $4.00/M output, and is Z.ai-direct only with no Foundry meter.

Pricing

per 1M tokens · USD / CHF
ModelTier / HostContextInputCachedOutputConfidence
GLM 5
The cheapest GLM lane on Foundry: ~29% below 5.1/5.2 on input and ~27% below on output, for the same 200K window as 5.1.
Foundry ·Data Zone
Fireworks-hosted
200K
128K max out
$1.10CHF 0.886$0.22CHF 0.177$3.52CHF 2.83official
Azure Retail Prices API 'FW GLM 5' meters, captured 2026-07-25 (effective 2026-06-01): input $0.0011/1K, output $0.00352/1K, cached input $0.00022/1K. Each is exactly 1.1x the Z.ai direct rate, the standard Data Zone premium. No Global-tier meter is published for GLM 5, matching 5.1 and 5.2.
GLM 5.1
Foundry ·Data Zone
Fireworks-hosted
200K$1.54CHF 1.24$0.286CHF 0.23$4.84CHF 3.90official
Fireworks official live pricing page.
GLM 5.2
Cached input is officially published at $0.15/M, ~48% below GLM 5.1's $0.286/M.
Foundry ·Data Zone
Fireworks-hosted
1M
131K max out
$1.54CHF 1.24$0.15CHF 0.121$4.84CHF 3.90official
Azure Retail Prices API 'FW GLM 5.2' meters, captured 2026-07-22 (effective 2026-07-01): input $0.00154/1K, output $0.00484/1K, cached input $0.00015/1K, uniform across regions. Earlier estimate (equal to GLM 5.1) was right on input/output but high on cache.
GLM-5
Z.ai publishes a cached-input rate for GLM-5 ($0.20/M). Cached-input storage is currently free for a limited time, a billing dimension this catalog does not model.
Direct API
Z.ai direct API
200K
128K max out
$1.00CHF 0.805$0.20CHF 0.161$3.20CHF 2.58official
Z.ai official pricing page, captured 2026-07-25: input $1, cached input $0.2, output $3.2 per 1M tokens. Model page lists a 200K window and 128K max output. The 'Limited-time Free' marker on that page applies only to the separate Cached Input Storage column, not to these rates.
GLM-5-Turbo
Agentic / tool-calling variant priced between GLM-5 and GLM-5.1/5.2; no Foundry meter, so Z.ai direct is the only lane.
Direct API
Z.ai direct API
200K
128K max out
$1.20CHF 0.966$0.24CHF 0.193$4.00CHF 3.22official
Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-07-29: $1.20/M input, $0.24/M cached input, $4.00/M output. The model page (docs.z.ai/guides/llm/glm-5-turbo) lists a 200K window and 128K max output and positions it as optimized for agentic long-chain execution and high-throughput tool calling. A full sweep of Azure's Foundry catalog on 2026-07-29 found no GLM-5-Turbo meter — only GLM 5, 5.1 and 5.2 are resold there.
GLM-5.1
Z.ai is the family's first-party API; input/output already matched Z.ai exactly, and the cache rate is now sourced from there too.
Direct API
Z.ai direct API
200K$1.40CHF 1.13$0.26CHF 0.209$4.40CHF 3.54official
Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-07-26: $1.40/M input, $0.26/M cached input, $4.40/M output. Vendor label corrected from 'Fireworks direct API' to 'Z.ai direct API' — Z.ai is the developer's own API per this catalog's Direct-tier definition, and previously had no cached-input figure recorded.
GLM-5.2
Identical direct rate to GLM-5.1. Z.ai is the family's first-party API.
Direct API
Z.ai direct API
1M
131K max out
$1.40CHF 1.13$0.26CHF 0.209$4.40CHF 3.54official
Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-07-26: $1.40/M input, $0.26/M cached input, $4.40/M output (5.1 and 5.2 priced identically). Vendor label corrected from 'Fireworks direct API' to 'Z.ai direct API', matching the GLM-5 Direct row and closing the family's mixed-vendor Direct lane.
GLM-5.3
Prices identically to GLM-5.1 and GLM-5.2 on all three dimensions. Keeps the family's 1M-token window with 128K max output, and always runs with reasoning enabled — three effort levels (low, high, max); disabling reasoning is not supported.
Direct API
Z.ai direct API
1M
128K max out
$1.40CHF 1.13$0.26CHF 0.209$4.40CHF 3.54official
Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-08-19: $1.40/M input, $0.26/M cached input, $4.40/M output per 1M tokens, USD — identical to GLM-5.1 and GLM-5.2. The model guide (docs.z.ai/guides/llm/glm-5.3) lists a 1M-token window, 128K max output, and reasoning always enabled with three effort levels (low, high, max). A full Azure Retail Prices API sweep on 2026-08-19 (29,405 rows) found no Foundry meter for GLM 5.3, so this direct-API row is the only lane.
GLM-5.3-Flash
Native multimodal GLM-5 model with image, video and file input. Reasoning is always enabled; the current 50% promotion covers input, cached input and output, then reverts to the published list rate on 2026-09-09 at 16:00 UTC.
Direct API
Z.ai direct API
1M
128K max out
$0.075CHF 0.06$0.015CHF 0.012$0.25CHF 0.201official
Z.ai official pricing page (docs.z.ai/guides/overview/pricing), captured 2026-08-28: GLM-5.3-Flash is listed at $0.075/M input, $0.015/M cached input and $0.25/M output, with strikethrough list prices of $0.15/$0.03/$0.50. The page states that the 50% promotion ends at 24:00 on September 9, 2026 (UTC+8, Singapore time). The official model guide (docs.z.ai/guides/vlm/glm-5.3-flash) lists image/video/text/file input, a 1M-token context window, 128K max output and reasoning that cannot be disabled. A full Azure Retail Prices API sweep on 2026-08-28 found no Foundry meter for GLM-5.3-Flash, so this direct-API row is the only lane.
List price (from September 10) begins 9 Sep, 16:00 UTC
List price (from September 10)$0.15 / $0.03 / $0.50
Confidenceofficial published pagederived reconciled from billing estimate pattern-inferred
Field notes

Quirks & gotchas

Insight

GLM 5's Data Zone rate is a clean 1.1x

All three of GLM 5's Foundry Data Zone meters land at exactly 1.1x the Z.ai direct rate: $1.00 to $1.10 input, $0.20 to $0.22 cached, $3.20 to $3.52 output. Two independently published sources agreeing on all three dimensions is the strongest confirmation this catalog gets that both numbers are right.

Insight

5.2 Data Zone: an estimate that held up

The unlisted GLM-5.2 Data Zone rate was set equal to GLM-5.1's ($1.54 / CHF 1.24 input, $0.286 / CHF 0.23 cached, $4.84 / CHF 3.90 output). Real invoices later matched almost exactly. A well-reasoned estimate, flagged as such, can hold until an official number lands.

Note

1M context window on 5.2

GLM-5.2's 1M-token window (up from 200K on 5.1) puts it in the same autonomy tier as DeepSeek V4 for long runs.

Other providers
KimiDeepSeekOpenAI / Azure OpenAIClaudeGeminiGrokQwenMistralMiniMaxEmbeddingsCompare all →