Skip to content
llm-spend
GitHub
History

Changelog

LLM pricing moves fast. This log records when a rate changed, a model launched, or the method changed. A reference is only trustworthy if it says how fresh it is.

  1. method

    Comparisons are now shareable, exportable, and keep a shortlist

    A comparison is no longer something you have to rebuild from memory. The workload, pricing scenario, filters, sort order, and pinned lanes are all written into the /compare address as you work, so copying that link hands someone the exact comparison you are looking at. “Copy scenario link” puts the full URL on the clipboard and, if the browser refuses, shows the link as selectable text instead. “Export CSV” writes the rows currently on screen — scenario label, provider, model, host, tier, active rate variant, resolved input/cache/output rates, workload total in USD and CHF, and confidence — straight to a file, entirely in the browser.

    Every result row now has a pin control. Pinned lanes collect in a tray that follows you down the page, survives a refresh, and expands into a side-by-side comparison of up to four lanes showing each one's resolved rates, blended input, and workload total, with the cheapest lane in the shortlist as the cost baseline and every other lane's difference in dollars and percent. That baseline is the lowest cost for the workload you selected — never a claim about model quality.

    Three new destinations open up alongside it. Each lane now has its own cost-anatomy page splitting a workload into fresh input, cached input, and output, with per-dimension provenance, the lane's full published rate schedule, same-model Direct-versus-Foundry markup, and cost-comparable alternatives. The new budget planner projects a representative month from a per-request workload and answers what cache hit rate a lane would need to fit a monthly budget. The new freshness page shows when the catalog was last audited in full, which lanes are missing a cache meter or carry a scheduled rate change, and where each provider's figures come from.

    No catalog price, cache rate, availability, confidence level, or calculation formula changed in this release. Every figure on the new pages is resolved by the same pricing code that already drives the comparison table.

  2. model

    Claude Fable 5.1 and Mythos 5.1 add cheaper cache reads

    Anthropic launched Claude Fable 5.1 on September 1 as its latest generally available frontier model. The catalog now includes a Direct row at $10/M input, $0.25/M cached input and $50/M output, with a 1M-token context window and 128K maximum output. Its cache-read price is one quarter of Claude Fable 5's $1/M rate; the published cache-write prices are unchanged.

    Invite-only Claude Mythos 5.1 is also listed as a limited-availability Direct row at the same $10/$0.25/$50 rates and specifications. Anthropic lists Fable 5.1 as available in Microsoft Foundry, but Foundry continues to bill Claude through Claude Consumption Units and the Azure Retail Prices API still publishes no Anthropic per-token meter, so no token-priced Foundry row is inferred.

  3. method

    Pinned lanes now follow you to the budget planner and lane pages

    The shortlist tray on /compare only ever showed up there — a lane pinned during a comparison became invisible the moment you moved to its own cost-anatomy page or to the budget planner. Both routes now show a compact shortlist strip naming every currently pinned lane by provider, model, and deployment tier, alongside the workload cost that page already has active for it (a lane's cost-anatomy page prices it under the same workload shown on /compare by default; the budget planner prices it under its own per-request shape, so the same pinned lane can read differently in each place, correctly). Each chip links to that lane's cost-anatomy page, and an unpin control keeps the shortlist in sync with the tray on /compare, since every surface reads and writes the same stored shortlist.

    The strip is entirely absent when nothing is pinned, and no catalog price, cache rate, availability, confidence level, or calculation formula changed in this release — every figure it shows is read from the same resolved comparison rows each page already computed for its own table or projection, never a second pricing pass.

  4. method

    The comparison table becomes a cost decision cockpit

    The cross-provider comparison now starts with four clearly illustrative workload shapes, then surfaces the lowest-cost lane overall, on Microsoft Foundry, and on a direct API before the full rate table. These are cost leaders for the selected workload and pricing scenario — not model-quality recommendations — and each signal remains traceable to the same sourced catalog rows below.

    Search, provider, deployment, cache-meter, and official-only filters now compose in one focused control surface. An explicit sort control stays synchronized with keyboard-sortable table headers, while a directional empty state preserves the workload and scenario if a filter combination finds nothing.

    The initial view shows the top 12 lanes under the active sort instead of forcing every visitor through all 66 rows; the complete filtered catalog remains one action away. On narrow screens those same semantic rows become legible cost cards with no document-level horizontal overflow. No catalog price, availability, confidence level, or calculation formula changed in this release.

  5. model

    GLM-5.3-Flash and Qwen3.8 Flash add new direct lanes

    Z.ai's new GLM-5.3-Flash is now priced in the catalog at its current 50%-off rate: $0.075/M input, $0.015/M cached input and $0.25/M output. The model is natively multimodal, keeps a 1M-token context window with 128K max output, and always runs with reasoning enabled. Z.ai states that the promotion ends at 24:00 on September 9, 2026 (UTC+8), so the row now switches to the published $0.15/$0.03/$0.50 list rate at 16:00 UTC on September 9.

    Alibaba's current Model Studio pricing adds Qwen3.8 Flash on the International Singapore endpoint at $0.15/M input and $0.47/M output for the full 1M-token band. Its first-party QwenCloud model page publishes $0.016/M for both implicit cached input and explicit cache reads, plus $0.20/M for explicit cache creation; the catalog uses the published $0.016 cached-input figure. Both new generation rows are direct-only: the full Azure Retail Prices sweep on 2026-08-28 found no matching Foundry meter.

    The same cache review corrected Qwen3.8 Max's cached-input cell from the old 10%-derived $0.20/M value to the official $0.17/M explicit cache-read rate. Alibaba now publishes $0.25/M for implicit cache, $2.50/M for explicit cache creation and $0.17/M for explicit cache reads, and explicitly calls Qwen3.8 models exceptions to the generic cache formula. Qwen3.7 Max's existing 50%-off row is also now documented as ending on August 31, 2026.

    Finally, Alibaba's Singapore Model Studio endpoint now lists Qwen3.7 text embedding at $0.07/M input, with up to 128,000 tokens per input line and no output or cached-input charge. It is added as an input-only direct embedding lane; no Foundry embedding meter was found.

  6. pricing

    DeepSeek's peak pricing is weekdays only — weekends bill at the off-peak rate

    DeepSeek's peak/off-peak split, which began billing on 2026-08-16, does not run every day. Its pricing page states that peak hours are 01:00-04:00 and 06:00-10:00 UTC "Monday through Friday (all other hours are off-peak)", and the Chinese edition of the same page agrees, expressing the identical windows in Beijing time. Every hour of Saturday and Sunday bills off-peak.

    This site was showing the peak rate during those hours at the weekend, which overstated the cost of a weekend DeepSeek call by exactly 2x — V4 Pro read $1.32/M input and $3.96/M output when the real weekend rate is $0.66 and $1.98, and V4 Flash read $0.44/$1.32 instead of $0.22/$0.66. Twelve hours of every week were priced wrongly. Both DeepSeek direct rows are now correct at every hour of every day, and the published rates themselves are unchanged — only the schedule they are applied on.

    Behind that, the catalog's scheduled-pricing model gains a day-of-week condition, so a rate that recurs only on working days can be described exactly rather than approximated. The rate the site quotes is still always a number DeepSeek publishes; what changed is that it is now applied on the calendar DeepSeek publishes it for.

  7. pricing

    OpenAI cuts GPT-5.6 Sol, Foundry does not follow — and DeepSeek V4 Flash gains two Data Zone lanes

    OpenAI cut its flagship on 2026-08-21. GPT-5.6 Sol now bills $4.00/M input, $0.40/M cached input and $20.00/M output on the direct API, down from $5.00/$0.50/$30.00, with the long-context tier at $8.00/$0.80/$30.00. OpenAI's pricing page labels this promotional and says it is "available at least through November 21, 2026" — an open-ended commitment with no published reversion rate, so nothing is staged here for a future date.

    Microsoft Foundry has not passed the cut on. A full sweep of the Azure Retail Prices catalog today found every GPT-5.6 Sol meter still on its original 2026-07-01 tranche at $5.00/$0.50/$30.00, with none of the newer 2026-08-01 tranche that carried the Terra and Luna cuts three weeks after those were announced. Running Sol on Foundry Global therefore costs 25% more per input token and 50% more per output token than going direct, and the Data Zone tier stacks its usual 10% on top. Terra and Luna are unaffected and remain at 1:1 parity. The Sol rates shown here are unchanged, because the Azure meter is what bills a Foundry customer — what changed is that it is no longer the same number OpenAI charges.

    Separately, DeepSeek V4 Flash gains the two Foundry Data Zone lanes that V4 Pro already had. Microsoft's first-party Data Zone deployment bills $0.21/$0.031/$0.56, the expected 10% over Global. The Fireworks-hosted Data Zone listing bills $0.15/$0.03/$0.31 — below the Global tier on every dimension, which inverts the usual rule that Data Zone is the premium tier. That makes it the cheapest DeepSeek lane tracked here and the cheapest input rate on the site. The two are different sellers rather than different tiers of one seller, so the tier label carries no price ordering across hosts.

  8. model

    Kimi K3 reaches Microsoft Foundry, at a 10% Data Zone premium

    Kimi K3 has been tracked here since July as a direct-API-only model: every Azure Retail Prices sweep since then found no Foundry meter for it, so there was no Foundry lane to compare against. That has changed. A Fireworks-hosted Data Zone listing went live with meters effective 2026-08-01, at $3.30/M input, $0.33/M cached input, and $16.50/M output.

    That is exactly 1.10x Moonshot's own $3.00/$0.30/$15.00 — the same Data Zone premium the Fireworks-hosted GLM, MiniMax, DeepSeek, and Kimi K2.6 lanes already carry. The listing is uniform across all 20 commercial Data Zone regions, with no regional outlier. There is no Foundry Global lane for K3, so unlike the K2 generation the only choice on Foundry is the Data Zone tier, and running K3 there costs 10% more than going direct to Moonshot.

    A caveat worth knowing if you go looking for these rates yourself: Microsoft separately published six K3 meters under its native Azure Kimi product whose names contradict their own prices — one meter tagged as output bills $0.33/M, another tagged as Data Zone output bills the Global input rate. Their prices do resolve into a plausible Global and Data Zone pair, but no dimension can be read off the meter names, so only the correctly-labelled Fireworks lane is quoted here.

  9. pricing

    Microsoft Foundry finally passes on the GPT-5.6 Terra and Luna price cuts

    OpenAI cut its direct-API rates for GPT-5.6 Terra and Luna on 2026-07-30, and for three weeks Microsoft Foundry kept billing the old ones. That gap is now closed. A Foundry meter tranche effective 2026-08-01 replaces the previous rates outright — the superseded 2026-07-01 meters no longer appear in the Azure retail catalog for either model — so all three GPT-5.6 variants are back at 1:1 parity with OpenAI direct.

    Terra Global drops 20%, from $2.50/$0.25/$15.00 to $2.00/$0.20/$12.00 per 1M tokens. Luna Global drops 80%, from $1.00/$0.10/$6.00 to $0.20/$0.02/$1.20, which makes it the cheapest lane tracked here. The long-context tiers move by the same proportions, to $4.00/$0.40/$18.00 and $0.40/$0.04/$1.80. The priority (Azure "PP", renamed "Fast mode" by OpenAI) meters were cut alongside the standard ones and still bill at exactly 2x, and the Data Zone meters followed too, at the usual 1.10x. GPT-5.6 Sol was never cut and is unchanged at $5.00/$0.50/$30.00 on both Foundry and direct.

    If you deferred a Foundry migration for these two models on cost grounds, the arithmetic has changed. It is also a reminder about sourcing: during the three-week lag, Azure support answers and at least one downstream cost tracker described the cut as already live on Azure while the retail catalog still billed the old rate. The published meter is what bills you.

  10. model

    GLM-5.3 is now tracked, with a published direct-API rate

    GLM-5.3, announced 2026-08-14 with no per-token price yet, now has one on Z.ai's pricing page: $1.40/M input, $0.26/M cached input, $4.40/M output — identical to GLM-5.1 and GLM-5.2 on all three dimensions. At the same price point, the choice within the family comes down to context window and reasoning behavior rather than cost: GLM-5.3 keeps the 1M-token window with 128K max output, and always runs with reasoning enabled (three effort levels — low, high, max — rather than an on/off toggle).

    It is direct-API only for now. A full Azure Retail Prices API sweep on 2026-08-19 found no Microsoft Foundry meter for GLM 5.3, so unlike GLM-5, 5.1, and 5.2 — each of which has a Fireworks-hosted Data Zone row on Foundry — there is no Foundry lane to compare it against yet.

  11. method

    Provider cards no longer advertise a DeepSeek rate that is no longer charged

    The homepage provider cards show a "from $X /M in" figure for each provider. That figure was computed from each row's base input rate without accounting for conditional pricing, so after DeepSeek's peak/off-peak billing took effect on 2026-08-16 16:00 UTC it advertised DeepSeek "from $0.14 /M in" — a rate that is no longer charged at any hour of the day.

    The figure is now resolved through the same rate engine the pricing table and compare page use, and reads $0.19 /M, the cheapest rate DeepSeek actually charges (its V4 Flash Global tier). No other provider's figure changed.

    Because the calculation is now variant-aware, it will also follow Qwen3.7 Max's return to list price on 2026-09-01 and the Gemini Flash reversions on 2027-01-01 automatically, instead of continuing to show a superseded rate.

  12. method

    Cohere Embed's direct rate is confirmed published again, and marked official

    Cohere's own pricing page publishes a per-token rate for Embed 4 again, at $0.12 per 1M tokens, shown under the "Advanced retrieval models" tab. The catalog's direct Cohere embedding row had been marked as a derived figure since 2026-07-29, when the rate appeared to have been removed from Cohere's site and Azure's meter became the only published source. The rate itself is unchanged — only its provenance label moves from derived to official, so the row no longer carries a confidence caveat it does not need.

    The earlier "removed" conclusion rested on archived snapshots, which capture only server-rendered HTML; this page renders its per-token rates in tabs that load client-side, so the rate could not have appeared in a snapshot whether or not it was published.

  13. method

    Compare page's Service tier picker now returns real prices for seven rows, not a no-op

    The compare / cost-calculator page's Service tier control (Standard / Batch / Flex / Priority / Highspeed), added in an earlier methodology update, has been structurally present but functionally inert: no catalog row defined a serviceTier-scoped rate variant, so picking anything but Standard silently degraded every row back to its base rate. Seven rows across four providers now carry real, officially published service-tier pricing (17 new variants total), so selecting a tier actually changes the numbers for these rows.

    Kimi K2.7 Code (Global) gains a Highspeed variant at exactly 2x standard on every dimension: $1.90/M input, $0.38/M cached, $8.00/M output versus the base $0.95/$0.19/$4.00 — Moonshot's own kimi-k2.7-code-highspeed listing. Gemini 3.6 Flash and Gemini 3.7 Flash (both Global) each gain Batch, Flex, and Priority variants, every one paired with the row's existing 2027-01-01 promo-reversion date so the right numbers show on both sides of that switch: Batch and Flex are both exactly 50% of Standard ($0.375/$0.0375/$1.875 now, $0.75/$0.075/$3.75 from 2027) and Priority is exactly 1.8x ($1.35/$0.135/$6.75 now, $2.70/$0.27/$13.50 from 2027) — Batch and Flex are numerically identical but billed as distinct service tiers, so both are modeled separately. MiniMax M3 (Direct) gains a Priority variant at exactly 1.5x standard: $0.45/M input, $0.09/M cached, $1.80/M output versus $0.30/$0.06/$1.20 — confirmed live against a distinct "Priority" tab on MiniMax's own pricing page; the same page shows no such tab for MiniMax M2.7, so no variant was added there. GPT-5.6 Sol, Terra, and Luna (all Foundry Global) each gain a Priority variant — Azure's own "PP" (priority processing) meters, exactly 2x each row's Standard Global rate: Sol $10.00/$1.00/$60.00, Terra $5.00/$0.50/$30.00, Luna $2.00/$0.20/$12.00. Azure's PP meters exist only for the short-context (ShortCo) listings; a full sweep found no matching Long Context PP meter for any GPT-5.6 variant.

    No base or Standard rate changed in this update, and picking "Standard" — the default — continues to show exactly what it showed before. This only makes the existing tier picker report real numbers for the rows above; every other row still correctly degrades to its base rate under a tier it doesn't publish.

  14. model

    DeepSeek V3 removed from the catalog

    Removed the standalone DeepSeek V3 row (Global tier, $1.14/M input, $4.56/M output, no cache meter) from the catalog. It's a pre-V3.1/V3.2/V4 generation model that no longer appears anywhere on DeepSeek's own pricing page — only deepseek-v4-flash and deepseek-v4-pro are listed today — and the V3.1 and V3.2 point releases that superseded it were already removed from this catalog previously.

    This does not affect the actively-tracked DeepSeek-V4 Pro and DeepSeek-V4 Flash rows (Direct, Global, and Data Zone lanes), which remain current and unchanged.

  15. method

    The compare page now prices every model under a chosen scenario, not just today's flat rate

    The compare / cost-calculator page gains scenario controls alongside the existing workload calculator: a Time picker (Now — live; Off-peak / Peak — representative hours; or any specific UTC hour) and a Service tier picker (standard / batch / flex / priority / highspeed). Every row's Input, Cached, Output, blended, and workload-cost figures — and the sort order — are now computed by resolving each model's rate under the selected scenario, via the same resolver added in the last two methodology updates, instead of always reading a row's flat base fields. The workload calculator's existing input-token control doubles as the prompt/context size fed into the resolver, since the compare page has no other notion of 'how big is this call' and no catalog row defines a context-band variant yet.

    This is what keeps the comparison honest once a scheduled rate change lands: today, picking any time or tier changes nothing, because DeepSeek's peak/off-peak split (2026-08-16T16:00:00Z), Qwen3.7 Max's promo reversion (2026-09-01), and Gemini 3.6/3.7 Flash's promo reversion (2027-01-01) haven't started yet — every row still resolves to its base rate, and the table, sort order and every displayed number are unchanged from before this update. From 2026-08-16 16:00 UTC onward, though, 'Now' will correctly show DeepSeek's actual peak or off-peak price instead of a stale flat rate, and picking 'Peak' or a specific UTC hour lets a visitor preview the other side of that split at any time. A small brand-colored label appears under a model's name whenever the selected scenario resolves it to a rate that differs from its base — restrained on purpose, since on most rows, most of the time, nothing differs.

    Time scenarios preview an hour of the day on the real current calendar date; they are not a time machine. Picking 'Peak' today does not make DeepSeek's not-yet-effective peak rate appear early, because the resolver's published effective date still gates it — the preview only overrides which hour is being asked about, never which day. No underlying rate changed in this update; it is comparison-tool methodology only, the same class of change as the two rate-variant updates that came before it.

  16. method

    Rate-variant rows now render the price actually in effect, not just the base rate

    The five rows carrying a scheduled rate variant — DeepSeek-V4 Pro (Direct), DeepSeek-V4 Flash (Direct), Gemini 3.6 Flash, Gemini 3.7 Flash, and Qwen3.7 Max (Promo) — now display whichever rate is actually in force right now, resolved with the same logic added in the previous methodology update, instead of always showing the row's flat base fields. Each also gets a small badge naming the active variant when one applies, a note on when the price next changes (a plain countdown once a variant regime is under way, or a '<label> begins <date>, UTC' announcement while it is still pending — none of today's five regimes has started yet), and every published variant's own numbers laid out underneath the row, so the full rate card is visible without clicking anything.

    No rate changed today. DeepSeek's peak/off-peak split still takes effect 2026-08-16T16:00:00Z, Qwen3.7 Max's promo still reverts 2026-09-01T00:00:00Z, and Gemini 3.6/3.7 Flash still revert 2027-01-01T00:00:00Z — this entry is the page catching up to data the catalog already had. The practical effect is that the DeepSeek rows stop showing a stale flat rate the moment the peak/off-peak switch lands: the displayed price now self-corrects in the visitor's browser, without waiting on a rebuild.

    All variant numbers are present in the server-rendered page, since they are static published facts; only which variant is marked 'active' and the countdown text depend on the reader's clock, and both refresh automatically every 30 seconds so a boundary crossed while the page is left open updates on its own.

  17. method

    DeepSeek, Gemini and Qwen's scheduled rate changes are now modeled as data

    No displayed rate changes today. This is a methodology update: the catalog can now record a rate that changes at a known future instant as structured data (an exact effective timestamp plus its own published numbers), instead of only describing the upcoming change in a row's notes text. Three rows that already carried this kind of prose description are the first to move onto the new structure.

    DeepSeek-V4 Pro (Direct) and DeepSeek-V4 Flash (Direct) each gain two scheduled variants, both effective 2026-08-16T16:00:00Z: a Peak rate for 01:00-04:00 and 06:00-10:00 UTC, and an Off-peak rate — exactly half of Peak on every dimension — for every other hour. Per 1M tokens (cache miss / cache hit / output): V4 Pro peak $1.32 / $0.044 / $3.96, off-peak $0.66 / $0.022 / $1.98; V4 Flash peak $0.44 / $0.014 / $1.32, off-peak $0.22 / $0.007 / $0.66. Today's flat rates ($0.435/$0.003625/$0.87 and $0.14/$0.0028/$0.28) are unaffected until that instant.

    Gemini 3.6 Flash and Gemini 3.7 Flash each gain one scheduled variant reverting to list price — $1.50/M input, $0.15/M cached, $7.50/M output — effective 2027-01-01T00:00:00Z, matching the reversion Google already publishes inline on its pricing page. Both rows keep showing today's promotional rate ($0.75/$0.075/$3.75) until then.

    Qwen3.7 Max (Promo) gains one scheduled variant reverting to list price — $2.50/M input, $0.25/M cached, $7.50/M output — effective 2026-09-01T00:00:00Z, the date Alibaba Cloud's campaign page states the discount 'runs until.' The row keeps showing today's discounted rate ($1.25/$0.125/$3.75) until then.

    The site still renders each row's flat rate directly, so none of this is visible on the pages yet — it is data plumbing for a later phase that will make the displayed price switch automatically at these instants.

  18. model

    Gemini 3.7 Flash added; Gemini 3.6 Flash's price halved through year-end

    Google's Gemini API pricing page now shows Gemini 3.6 Flash at $0.75/M input, $0.075/M cached input, and $3.75/M output — half its previous $1.50/$0.15/$7.50 rate. The page publishes this as a promotional rate running through December 31, 2026, reverting to $1.50/$0.15/$7.50 on January 1, 2027; each price cell states this inline ("$0.75 through December 31, 2026. $1.50 starting January 1, 2027.").

    A new Gemini 3.7 Flash — Google's own description: "our most capable Flash model for agentic workflows and multimodal reasoning" — is added to the catalog as the new flagship, tracked at the identical promotional rate: $0.75/M input, $0.075/M cached, $3.75/M output, with the same reversion terms on 2027-01-01.

    Gemini 3.5 Flash is unaffected and stays listed for comparison at $1.50/M input, $0.15/M cached, $9.00/M output. Google also publishes Batch and Flex pricing at 50% of the standard rate and a Priority tier at 1.8x, but this site's schema does not model those tiers.

  19. pricing

    DeepSeek's price increase gets a date: peak/off-peak billing from August 16

    DeepSeek's pricing page footnote — previously undated — now gives an exact effective time for the API's long-signalled price increase: peak / off-peak billing begins at 16:00 UTC on August 16, 2026. Off-peak rates are set at exactly half of peak. Peak hours are 01:00–04:00 and 06:00–10:00 UTC; every other hour is off-peak.

    Current flat rates ($0.435/M input, $0.003625/M cached, $0.87/M output for V4 Pro; $0.14/M input, $0.0028/M cached, $0.28/M output for V4 Flash) remain in force until that moment, and the catalog's two DeepSeek Direct rows are unchanged for now.

    From 2026-08-16 16:00 UTC, the new rates (per 1M tokens, cache hit / cache miss / output) are: DeepSeek-V4 Flash off-peak $0.007 / $0.22 / $0.66, peak $0.014 / $0.44 / $1.32. DeepSeek-V4 Pro off-peak $0.022 / $0.66 / $1.98, peak $0.044 / $1.32 / $3.96.

    Even the off-peak rate is an increase on every dimension versus today's flat rate — e.g. V4 Pro output rises 2.28x ($0.87 to $1.98), V4 Flash output rises 2.36x ($0.28 to $0.66), and V4 Pro's cache-hit rate rises from $0.003625/M to $0.022/M. Peak-hour rates run roughly double the off-peak figures again.

    This affects only DeepSeek's first-party direct API. The Microsoft Foundry (Global/DataZone) and Fireworks-hosted DeepSeek rows are Azure/Fireworks meters and are unaffected. This site's schema does not model time-of-day pricing, so the change taking effect will require a scope decision on how to represent it.

  20. method

    Mistral Medium 3.5's cached input rate is now officially published

    Mistral now publishes Mistral Medium 3.5's cached-input rate as an explicit dollar figure — $0.15/M — on its new consolidated inference pricing page, alongside Input ($1.50/M) and Output ($7.50/M). No rate changed: this is the same $0.15/M figure the catalog already carried, previously back-derived from Mistral's published -90%-off cache discount rule rather than read directly off a price list. The Direct row's cached-input confidence is upgraded from derived to official to reflect the new first-party source, which supersedes the previously-cited mistral.ai/pricing/api page.

    The two Mistral Foundry rows (Global and DataZone) are unchanged and still correctly show no cached rate — no Azure cache meter exists for Mistral Medium 3.5 on either tier, re-confirmed in today's full Foundry sweep.

  21. model

    Grok 4.6 added — xAI's new flagship, direct API only

    xAI released Grok 4.6 today, its new flagship model, labeled "Latest" in xAI's docs and superseding Grok 4.5 as the headline model. On xAI's direct API it prices at $2.00/M input, $0.50/M cached input, and $6.00/M output for prompts under 200K tokens, doubling to $4.00/M, $1.00/M, and $12.00/M above that threshold — the same long-context doubling pattern as Grok 4.5. Input and output prices are identical to Grok 4.5's, but cached input is higher ($0.50/M vs. $0.30/M).

    Grok 4.6 has a 500K-token context window and accepts text and image input with text output. It is direct-API only: a full Azure Retail Prices API sweep on 2026-08-12 found zero meters referencing "4.6", so it is not yet available on Microsoft Foundry. The existing Grok 4.5 row is unchanged and stays listed for comparison ($2.00/$0.30/$6.00, re-verified today).

    Separately, Anthropic's pricing docs page has caught up to the permanent-pricing announcement covered in the 2026-08-11 entry below: it now shows a single Claude Sonnet 5 row and states verbatim that the $2/$10 per-million rate "is now the standard price" and the planned September 1 increase "will not occur." This closes the docs-page-lag caveat that had been carried in both Sonnet 5 sourceNotes.

  22. pricing

    Claude Sonnet 5's introductory pricing made permanent — planned Sept 1 increase cancelled

    Anthropic's official @claudeai account announced on X at 9:03pm on 2026-08-10 that Claude Sonnet 5's introductory pricing — $2/M input, $0.20/M cached input, $10/M output — is now permanent: "We launched Sonnet 5 in June at $2 per million input tokens and $10 per million output tokens through August 31, and that price will remain unchanged." This reverses the previously-published plan for the rate to rise to $3/M input, $0.30/M cached, $15/M output on 2026-09-01.

    The catalog's separate "Standard" rows (Direct and the Foundry CCU estimate) have been removed, since that $3/$0.30/$15 rate will now never take effect. The two "Intro" rows are renamed to plain "Claude Sonnet 5" — the rate was never actually time-limited after all, it just stopped rising — and their notes/sourceNote now record the cancellation. Rates are unchanged: $2/M input, $0.20/M cached, $10/M output on both the Direct row and the Foundry (Global, estimate) row.

    Anthropic's own pricing page had not been updated as of 2026-08-11 and still shows the old two-tier structure — "$2/$10 through August 31, 2026" and "$3/$15 starting September 1, 2026" as separate rows. Treat that as the page not having caught up to the X announcement yet, not as a contradiction; the X post is the authoritative current source until the docs page is revised.

  23. pricing

    Qwen3.7 Max's 50% discount now has an official end date: August 31, 2026

    Alibaba Cloud's campaign page for Qwen3.8-Max ("Qwen3.8-Max is Here") states, in two places, that the Qwen3.7-Max limited-time 50% discount "runs until August 31, 2026," and that it "applies to all 4 billing items: Input, Output, Explicit Cache Creation, and Explicit Cache Hit." This is the first official end date published for a promo that has been open-ended since it launched — Model Studio's own pricing page still shows only the undated "Limited-time 50% off" label.

    No rate change today: the tracked promo rate ($1.25/M input, $0.125/M cached, $3.75/M output) stays in effect through 2026-08-31. From 2026-09-01 it reverts to list price ($2.50/M input, $0.25/M cached, $7.50/M output) — the same day Claude Sonnet 5's introductory pricing ends.

    Qwen3.7 Plus's separate 20%-off discount is unaffected by this and still carries no published end date anywhere.

  24. method

    Provenance fix: OpenAI embedding rows now cite their Foundry meter, not OpenAI's pricing page

    The three OpenAI embedding rows (text-embedding-3-large, text-embedding-3-small, text-embedding-ada-002) are Foundry Global-tier entries, but each carried a sourceNote reading only "OpenAI pricing page." That citation was wrong on two counts: a Foundry Global row should cite the Azure meter that actually publishes the rate, and OpenAI's own pricing page no longer lists per-token embeddings pricing at all — confirmed by a raw DOM dump (15 pricing tables, none with an embeddings row; "ada-002" appears 0 times) and by Internet Archive snapshots from 2026-06-05 through 2026-08-04, none of which show an embeddings row either.

    Each sourceNote now cites the live Azure Retail Prices API meter instead: 'text-embedding-3-large-glbl Tokens' at $0.00013/1K ($0.13/M) across 17 commercial regions, 'text-embedding-3-small-glbl Tokens' at $0.00002/1K ($0.02/M) across 17 regions, and 'embedding-ada-glbl Tokens' at $0.0001/1K ($0.10/M) across 15 regions — all effective 2024-06-01, captured 2026-08-07. The figures are independently corroborated by OpenAI's embeddings guide, which publishes pages-per-dollar at roughly 800 tokens/page (62,500 for -3-small, 9,615 for -3-large, 12,500 for ada-002) that invert to exactly these three rates.

    Rates are unchanged — still $0.13 / $0.02 / $0.10 per M input, no output or cached price — and confidence stays official on all three rows, the same treatment already given the Cohere Embed v4 Foundry rows: the Azure meter is itself an official published per-token rate. Only the provenance shown on the site changed.

  25. model

    Qwen3.8 Max added — Alibaba's new flagship gets a per-token price

    Alibaba's new flagship qwen3.8-max now has published per-token pricing on Model Studio's International endpoint: $2.00/M input and $6.00/M output, a single price tier across the full 0–1M token window with Non-Thinking and Thinking modes priced identically. It is GA, not a preview label. This closes a watch item open since 2026-07-23 — until today the model was reachable only via Token Plan credits, with no published per-token rate.

    Cached input is derived at 10% of input ($0.20/M) per Model Studio's published context-cache rule, the same convention used across the rest of the Qwen family. The model also carries a 1M-token free quota valid for 90 days. A separate Global deployment scope prices the model lower, at $1.65/M input and $4.951/M output; the tracked lane stays International, matching every other Qwen row in this catalog.

    Qwen3.7 Max and Plus's limited-time promotional discounts remain live, and qwen3.6-max-preview's 2026-10-10 deprecation (with qwen3.7-max as the named replacement) is unaffected by this addition.

  26. pricing

    OpenAI cuts GPT-5.6 Terra and Luna direct-API prices; Microsoft Foundry has not followed

    OpenAI's official changelog confirms a direct-API price cut effective 2026-07-30: GPT-5.6 Terra costs 20% less (now $2.00/M input, $0.20/M cached, $12.00/M output; long context $4.00/$0.40/$18.00), and GPT-5.6 Luna costs 80% less (now $0.20/M input, $0.02/M cached, $1.20/M output; long context $0.40/$0.04/$1.80). Sol is unchanged at $5.00/$0.50/$30.00.

    A full Azure Retail Prices API sweep run today found the Foundry serverless meters for both models unchanged: Terra Global is still $2.50/$0.25/$15.00 (long context $5.00/$0.50/$22.50) and Luna Global is still $1.00/$0.10/$6.00 (long context $2.00/$0.20/$9.00), both still effective 2026-07-01. That breaks the Foundry-matches-direct 1:1 parity that had held since 2026-07-21: Azure now runs about 1.25x direct on Terra and about 5x direct on Luna. Sol's parity is unaffected since its direct price didn't move.

    No catalog rate numbers changed as part of this entry — the Foundry rows still reflect the live Foundry meters. The four affected rows' notes and sourceNotes were updated to record the direct-API cut and flag the gap; watch for the Foundry meters to follow.

  27. model

    Qwen3.6 Max Preview receives an official deprecation date

    Alibaba Model Studio now schedules the tracked qwen3.6-max-preview model for deprecation on October 10, 2026, with qwen3.7-max as the documented replacement. The catalog keeps the existing preview row for the time being and flags the deadline rather than removing a still-available model early.

    The current International pricing is unchanged: $1.30/M input and $7.80/M output for the ≤128K band tracked here (128K–256K remains $2/$12). No new Microsoft Foundry serverless token meter was found for Qwen in the 2026-08-01 sweep.

  28. pricing

    Cohere Embed v4 gains officially metered Foundry lanes; GLM-5-Turbo and MiniMax's current lineup added

    Cohere Embed v4 is now shown on Microsoft Foundry Global ($0.12/M) and Data Zone ($0.132/M), both officially metered. Cohere itself no longer publishes a per-token Embed rate — its own pricing page now shows only dedicated Model Vault instances — so the direct-lane figure is marked derived rather than official; Azure's meter is currently the only published source for the $0.12/M rate.

    Added GLM-5-Turbo at $1.20/M input, $0.24/M cached input, $4.00/M output — priced between GLM-5 and GLM-5.1/5.2, with a 200K context window. It is Z.ai direct only; no Foundry meter exists for it.

    Added MiniMax M3 and MiniMax M2.7 as new direct-API lanes, both at $0.30/M input, $0.06/M cached input, $1.20/M output. MiniMax M2.5, already tracked on Foundry, is now listed as a legacy model on MiniMax's own pricing page. M3's rate carries a 'Permanent 50% off' label with the list price struck through at exactly 2x and no stated end date.

    A full Foundry sweep on 2026-07-29 found no price change to any of the roughly 43 already-tracked rows, and no meter with an effective date on or after 2026-08-01.

  29. method

    Provenance fixes: Mistral Medium 3.5 Global effective date, Grok-4.3 Data Zone long-context rates documented

    Corrected the Mistral Medium 3.5 Foundry Global row's sourceNote: the 'MM3.5 Inp/Outp glbl' meters are effective 2026-06-01 across 44 regions (including all 9 APAC regions), not 2026-07-01 as previously stated — the only 2026-07-01 Global row is the malaysiawest region onboarding at the same price. No prices changed; this is a documentation-only correction.

    Documented the Grok-4.3 Data Zone long-context rates, which exist alongside the already-tracked Global long-context note but had no Data Zone equivalent: $2.75/M input, $0.44/M cached input, $5.50/M output — a clean 1.1x the Global long-context rate ($2.50/$0.40/$5.00), consistent with the headline Data Zone premium. Headline Grok-4.3 Data Zone prices ($1.375/$0.22/$2.75) are unchanged.

    A full sweep of the Azure Retail Prices API found no new models and no price changes: Grok 4.5, any Qwen serverless per-token meter, and any Claude/Anthropic meter are all still absent from Foundry, and no meter carries an effectiveStartDate on or after 2026-08-01 — the anticipated non-Global price increase has not yet appeared.

  30. pricing

    APAC Data Zone carries a 1.20x premium, not 1.10x, on first-party OpenAI lines

    A full sweep of the Azure Retail Prices API turned up a region split the catalog's single Data Zone row per model had been hiding: for Microsoft's first-party OpenAI lines, Data Zone is two prices, not one. US/EU data-zone regions bill the already-tracked 1.10x Global premium, but APAC data-zone regions (australiaeast, centralindia, eastasia, japaneast, japanwest, jioindiawest, koreacentral, southeastasia, southindia) bill a separate 1.20x Global premium, effective 2026-06-01. Concretely: GPT-5.2 Data Zone runs $1.925/$0.1925/$15.40 per M in US/EU but $2.10/$0.21/$16.80 in APAC; GPT-5.5 runs $5.50/$0.55/$33.00 in US/EU but $6.00/$0.60/$36.00 in APAC (long-context and the text-embedding-3-large/small rows scale the same way). GPT-5.3 chat and the whole GPT-5.6 family (Sol/Terra/Luna) have no APAC data-zone rows yet, and no third-party family (Grok, Kimi, GLM, MiniMax, Mistral, DeepSeek) shows this split at all — it's exclusive to Microsoft's first-party OpenAI-hosted lines. A customer deploying in an APAC region should budget for a real 20% Foundry premium, not the 10% a single Data Zone number implies.

    Added a new native DeepSeek-V4 Pro Data Zone lane at $1.91/M input, $0.16/M cached, $3.83/M output — a first-party Foundry deployment distinct from the Fireworks-hosted Data Zone lane already listed, and slightly cheaper than it ($1.925/$0.165/$3.828).

    Everything else re-verified unchanged: a full sweep of the Foundry price feed plus the direct-API pricing pages for every tracked provider found no rate changes. Grok 4.5, any Qwen serverless per-token meter, and any Claude/Anthropic meter still have no Foundry listing at all, and no meter in the feed carries an effective date on or after 2026-08-01 — the anticipated non-Global price increase has not yet appeared. Qwen3.7 Max and Plus's promotional discounts are still live, and Claude Sonnet 5's introductory pricing still runs through 2026-08-31.

  31. pricing

    MiniMax added, new Qwen and Grok/GPT-5.6 Foundry lanes, GLM and Mistral cache rates, corrected GPT-5.2/5.3 cache figures

    Added a new MiniMax provider lane: MiniMax M2.5 and MiniMax 3 are resold on Microsoft Foundry as Data Zone-only serverless listings ($0.33/M input, $1.32/M output on both; cached input $0.033/M for M2.5 and $0.066/M for MiniMax 3), the same Data Zone-only situation as the GLM family — no Global-tier meter is published for either model.

    Added Qwen3.7 Flash, a new cheapest Qwen lane, at $0.10/M input and $0.40/M output (the 32K–256K context band, chosen to match the band already tracked for the existing Qwen3.6 Flash row so the two are directly comparable; the full published ladder runs $0.03/$0.13 at ≤32K up to $0.20/$0.80 at 256K–1M). Cached input is derived at 10% of input ($0.01/M), the same convention used across the rest of the Qwen family.

    Added two new official Foundry lanes for already-tracked models: Grok-4.3 Data Zone ($1.375/M input, $0.22/M cached, $2.75/M output — a clean 1.1x the tracked Global rate), and GPT-5.6 Terra Long Context ($5.00/$0.50/$22.50) and GPT-5.6 Luna Long Context ($2.00/$0.20/$9.00) on the Global tier, both previously only mentioned in a sibling row's notes.

    GLM-5.1 and GLM-5.2's Direct-lane rows now carry Z.ai's officially published cached-input rate ($0.26/M) and are relabeled 'Z.ai direct API' (from 'Fireworks direct API'), making the family's Direct lane internally consistent — Z.ai is the developer's own first-party API and its input/output rates already matched exactly. Mistral Medium 3.5's direct-API row now carries a derived cached-input rate ($0.15/M) based on Mistral's published -90%-off-input cache rule; both Foundry tiers remain uncached, confirmed against the Azure Retail Prices API.

    Corrected three rounded cached/input figures to their exact Azure meter values: GPT-5.3 Codex/Chat and GPT-5.2/Codex (Global) cached input from $0.18/M to $0.175/M, and GPT-5.2/Codex (Data Zone) input from $1.93/M to $1.925/M and cached input from $0.20/M to $0.1925/M.

  32. model

    Claude Opus 5 added — new Anthropic flagship, same pricing as Opus 4.8

    Added Claude Opus 5, now GA per Anthropic's official Claude API pricing page. It prices identically to Opus 4.8 at $5/M base input, $0.50/M cache-hit input, and $25/M output, with the same 1M-token context window — no launch premium.

    Opus 5 uses the same newer tokenizer introduced with Claude 4.7-and-later models (and Claude Mythos Preview), which can use roughly 30% more tokens for the same text versus pre-4.7 models. As with Opus 4.8, there is no separate Foundry/Global listing for Opus 5 in this catalog yet — only the direct-API rate is tracked.

  33. model

    GLM-5 added — the cheapest lane in the GLM family

    Added the original GLM-5, which had been missing from the catalog even though it is still generally available and undercuts both GLM-5.1 and GLM-5.2. On Microsoft Foundry's Data Zone tier it bills $1.10/M input, $0.22/M cached input, and $3.52/M output — about 29% below 5.1/5.2 on input and 27% below on output, with the same 200K context window as 5.1 and 128K max output.

    Z.ai's own API lists it at $1.00/M input, $0.20/M cached input, and $3.20/M output. Every Foundry Data Zone meter is exactly 1.1x the corresponding direct rate, the standard Data Zone premium, so both sets of numbers corroborate each other on all three dimensions. As with GLM-5.1 and 5.2, no Global-tier meter is published — Data Zone is the only Foundry tier for this family.

    If you do not need GLM-5.2's 1M-token window, GLM-5 is the value pick in this family rather than a superseded model. Note that Z.ai's 'Limited-time Free' marker applies to cached-input storage, a per-hour billing dimension this catalog does not model, and not to the per-token rates above.

  34. model

    Mistral Medium 3.5 added — new Mistral lane on Foundry

    Added a Mistral provider lane with Mistral Medium 3.5. Under the Microsoft–Mistral partnership announced 2026-07-21, the model is now resold on Microsoft Foundry as a serverless listing. Azure's Retail Prices API publishes 'MM3.5' meters at $1.50/M input and $7.50/M output on the Global tier (effective 2026-07-01) — identical to Mistral's own direct API rate for mistral-medium-latest.

    Foundry also publishes a Data Zone tier at $1.65/$8.25 per M (a clean 1.1x premium, effective 2026-06-01). No cached-input meter exists on any tier yet, so repeated-prompt workloads get no cache discount on this model. Mistral OCR 4, announced alongside it, bills per page rather than per token and is out of this catalog's scope.

  35. pricing

    Foundry cached-input meters now official for Kimi K2.5, K2.6, and Grok-4.3

    Azure's Retail Prices API now publishes dedicated cached-input meters (effective 2026-07-01 for Kimi, 2026-05-01 for Grok) that upgrade three catalog rows to official:

    Kimi K2.5 Thinking (Global): cached input added at $0.10/M — previously untracked.

    Kimi K2.6 (Global): cached input corrected to $0.16/M official, replacing the ~$0.19/M estimate that had been back-solved from a billing export.

    Grok-4.3 (Global): cached input added at $0.20/M — previously listed as having no cache meter. The meter is named bare '4.3' in the retail API (a 'grok' search misses it) and matches xAI's direct Grok 4.3 cache rate; long-context (≥200K) requests bill cached input at 2x ($0.40/M).

  36. model

    Gemini 3.6 Flash added; cached input now published for Flash line

    Added Gemini 3.6 Flash at $1.50/M input, $7.50/M output — the same input price as 3.5 Flash with ~17% cheaper output (50% batch discount: $0.75/$3.75). It's now the catalog's headline Gemini pick for coding and agentic workloads.

    Google's pricing page now also publishes a cached-input rate of $0.15/M for both 3.6 and 3.5 Flash, previously untracked for 3.5. Cache storage bills separately at $1.00 per 1M tokens per hour, a dimension noted in the source but not modeled by this site's schema.

  37. pricing

    GLM 5.2 Data Zone cached input corrected to official $0.15/M

    Azure now publishes dedicated 'FW GLM 5.2' Foundry meters (effective 2026-07-01), separate from GLM 5.1. Input and output are confirmed unchanged at $1.54/M and $4.84/M, but cached input is officially $0.15/M — well below the $0.286/M estimate the catalog had been carrying (inherited from GLM 5.1, since Fireworks previously charged both versions identically).

    GLM 5.2 Data Zone upgraded from estimate to official across the board.

  38. pricing

    GPT-5.6 Azure rates confirmed official; Sol Data Zone and Long Context added

    Azure's Retail Prices API now publishes Foundry meters for the whole GPT-5.6 family (effective 2026-07-01), confirming the 1:1 pattern: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per M tokens on Global, cached input at 10% of input, cache writes at 1.25x. All three entries upgraded from estimate to official.

    Added two Sol rows now that the meters are public: Data Zone at $5.50/$0.55/$33 (~10% premium over Global) and Long Context Global at $10/$1/$45. Terra and Luna have matching Data Zone and long-context meters, noted on their entries.

    Checked with no change needed: Grok 4.5 still has no Foundry meter (the retail catalog tops out at the Grok-4.x meters already tracked), and Qwen still has no serverless per-token Foundry meter (only Qwen3 32B fine-tuning hosting meters exist).

    Corrected Grok 4.5 cached input from the earlier ~$0.50/M third-party estimate to xAI's now-published official rate of $0.30/M.

  39. model

    Qwen3.7 flagships added; Alibaba cache rule confirmed official

    Added Alibaba's current flagship family, missed in yesterday's Qwen capture: Qwen3.7 Max at $1.25/M input, $3.75/M output effective (list $2.50/$7.50 under a limited-time 50% discount with no published end date; single price tier across a 1M-token window) and Qwen3.7 Plus at $0.32/$1.28 effective ≤256K, $0.96/$3.84 beyond (list rates under a 20% limited-time discount). Both priced on Model Studio's International (Singapore) endpoint.

    Alibaba's Context Cache doc now officially states the billing rule — explicit cache hits at 10% of input, cache creation at 125%, implicit hits at 20% — and lists every catalog Qwen model as supported. All Qwen cached-input rates upgraded from estimate to derived (10% of the billed input rate).

    Checked with no change needed: Grok 4.5 is still absent from Microsoft Foundry (the Grok tab still tops out at Grok-4.3 Global, $1.25/$2.50, no cache meters), Qwen still has no serverless per-token Foundry listing, and GPT-5.6 is still not on Azure's public OpenAI pricing page.

  40. model

    Grok (xAI) and Qwen (Alibaba) added to the catalog

    Added Grok: the new flagship Grok 4.5 at xAI's direct rates ($2/M input, $6/M output, 500K context; all meters double at ≥200K prompt tokens). Cached input was initially catalogued at ~$0.50/M from third-party listings and flagged estimate; xAI later published a $0.30/M official rate, recorded in the 2026-07-21 entry above. Grok 4.5 is not on Microsoft Foundry yet — the Foundry lineup tops out at Grok-4.3 Global ($1.25/$2.50) and Grok 4.1 Fast ($0.20/$0.50), and no Foundry Grok listing publishes a cached-input meter.

    Added Qwen: Alibaba Model Studio International rates for Qwen3.6 Plus ($0.50/$3.00 up to 256K, $2/$6 beyond), Qwen3.6 Flash ($0.25/$1.50), and Qwen3.6 Max Preview ($1.30/$7.80). Cache hits bill at ~10% of input per the published context-cache rule (marked estimate pending per-model confirmation). On Foundry, Qwen is Managed Compute only (GPU-hour, $4-8/hr) — there is no serverless per-token Qwen meter to compare.

  41. model

    Kimi K3 added; outdated catalog entries pruned

    Added Kimi K3's direct API pricing: $3/M cache-miss input, $0.30/M cache-hit input, and $15/M output. K3 is available now with a 1M-token context window.

    Removed older GPT-5, DeepSeek, Kimi Thinking, Gemini 3 Pro, and Claude Opus 4.8 Fast Mode entries to keep the comparison catalog focused on relevant current options.

  42. pricing

    DeepSeek direct cache rates and Azure meter provenance corrected

    Added DeepSeek's first-party V4 Pro ($0.003625/M) and V4 Flash ($0.0028/M) cache-hit rates, so the calculator now applies the selected cache-hit rate to direct API estimates.

    Upgraded Azure Global V4 Pro ($0.145/M) and V4 Flash ($0.028/M) cached-input rates from derived to official after confirming the exact meters in Azure's public retail catalog. Billing exports reconcile to those published rates.

  43. pricing

    Claude cache pricing: multiplier → explicit published rates

    Anthropic's pricing page now publishes per-model cache hit rates (e.g., Opus 4.8 cache read = $0.50/MTok, Sonnet 5 = $0.20/MTok intro / $0.30 standard, Fable 5 / Mythos 5 = $1.00/MTok), replacing the earlier multiplier approximation. All Claude entries updated with exact `cachedUsd` values.

  44. model

    Claude Fable 5 and Mythos 5 added

    Added Claude Fable 5 ($10 / $50 per MTok input/output) and Claude Mythos 5 ($10 / $50, limited availability) — Anthropic's next-gen frontier models — matching their published direct pricing. Same cache model as other Claude models: reads ~10% of input, writes ~1.25x input.

  45. pricing

    Heads-up: non-Global Foundry prices rising 2026-09-01

    Per the Azure Foundry Models pricing page, EU Data Zone and other non-US Regional deployment prices are set to increase on 2026-09-01. Global deployments are unchanged. No specific increase amount is published yet, so budget for a change and re-check closer to the date.

  46. model

    Microsoft Foundry rebrand; Claude now GA and Azure-hosted

    Microsoft renamed Azure AI Foundry to Microsoft Foundry. The site now uses the new name throughout; older entries below keep their original wording as a historical record.

    Claude Opus 4.8, Sonnet 5, and Haiku 4.5 are now GA and natively hosted on Microsoft Foundry (Azure-hosted, not just resold), billed through Azure via Claude Consumption Units (CCU) instead of the old per-model Azure meters. Added Sonnet 5 Foundry pricing, mirroring the direct rates: $2 / $10 per M input/output through 2026-08-31, then $3 / $15 per M from 2026-09-01. Microsoft doesn't publish a separate CCU-to-dollar ratio or an independent Foundry-native price, so the effective $/M is inherited from Anthropic's direct rate.

  47. launch

    Site launch

    llm-spend goes live: Kimi, DeepSeek, GLM, OpenAI / Azure OpenAI, Claude, Gemini, and embeddings, all in USD and CHF (reference 1 USD ≈ 0.805 CHF). Prices measured mostly on Azure AI Foundry, plus direct APIs.

    Includes the sticker-price methodology, the cache-economics case study, a workload cost calculator, and the RPM-vs-TPM explainer.

  48. model

    GPT-5.6 (Sol / Terra / Luna) reaches GA

    OpenAI's GPT-5.6 family hit GA: Sol (flagship), Terra (balanced), Luna (fast / cheap). Not on Azure's public page yet, so listed at OpenAI's direct rates as a high-confidence estimate for Azure via the 1:1 pattern.

    GPT-5.6 also bills cache writes at 1.25x the uncached input rate (was the standard input rate); reads stay ~90% off.

  49. pricing

    DeepSeek cuts direct API pricing 75%

    DeepSeek's direct API dropped V4 Pro to $0.435 / CHF 0.35 input, $0.87 / CHF 0.70 output, and V4 Flash to $0.14 / CHF 0.11 input, $0.28 / CHF 0.23 output. That widened the gap to cloud "Global" tiers, one reported at ~4.5x the direct price.

  50. method

    Hidden cache meters reconciled from billing exports

    Internal Azure Cost Management exports grouped by meter revealed undocumented cached-input meters on DeepSeek V4 Pro / V4 Flash "Global" and Kimi K2.7 in Azure AI Foundry. Back-solving against known input/output rates gave derived cache rates (~$0.145/M for V4 Pro, ~$0.028/M for V4 Flash, ~$0.19/M for Kimi). The links below document the public export and pricing methodology; the customer billing export itself is private. Rates were flagged derived pending official publication.