Skip to content
llm-spend
GitHub

Qwen

Alibaba

Qwen3.8 Max is the flagship, while Qwen3.8 Flash adds a new $0.15/$0.47 multimodal direct lane — but Foundry still has no per-token Qwen meter.

Qwen3.8 Max is Alibaba's flagship, GA with a plain $2/M input, $6/M output rate — no promotional discount. Qwen3.8 Flash adds a new multimodal 1M-context lane at $0.15/M input and $0.47/M output, while Qwen3.7 Max remains available at a limited-time 50%-off rate ($1.25/$3.75 effective) through 2026-08-31, alongside Qwen3.7 Plus at 20% off. All prices are Alibaba Cloud Model Studio's International (Singapore) endpoint.

Qwen3.6 Max Preview is scheduled for deprecation on 2026-10-10, with Qwen3.7 Max named as its replacement. On Microsoft Foundry, Qwen models are available only as Managed Compute — dedicated GPU-hour billing ($4–8 per compute hour) with no serverless per-token listing, so there is no Foundry token rate to compare.

Pricing

per 1M tokens · USD / CHF
ModelTier / HostContextInputCachedOutputConfidence
Qwen3.8 Max
New flagship, GA (not preview); text plus image/video understanding. Single price tier across the full 1M window; thinking and non-thinking modes priced the same. Plain rate — no promotional discount, unlike Qwen3.7 Max. The cached column uses the official $0.17/M explicit cache-read rate; implicit cache is $0.25/M. 1M-token free quota for 90 days.
Direct API
Model Studio (Intl)
1M$2.00CHF 1.61$0.17CHF 0.137$6.00CHF 4.83official
Alibaba Cloud Model Studio's qwen3.8-max model page, Singapore/International endpoint, captured 2026-08-28: $2/M input, $6/M output, $0.25/M implicit cached input, $2.50/M explicit cache creation and $0.17/M explicit cache read. The cached column uses the published explicit cache-read rate, matching the meaning used by the other Qwen cached-input rows; the earlier 10% derivation was not valid for this Qwen3.8 exception. A separate Global deployment-scope table prices this model lower, at $1.65/M input and $4.951/M output; the tracked lane is International.
Qwen3.8 Flash
New hybrid-thinking multimodal model with image, video and text input, a 1M-token context window and 131K max output. The International rate has no promotional label; its implicit-cache and explicit-cache-read rates are both $0.016/M.
Direct API
Model Studio (Intl)
1M
131K max out
$0.15CHF 0.121$0.016CHF 0.013$0.47CHF 0.378official
Alibaba Cloud Model Studio pricing page, Singapore/International endpoint, captured 2026-08-28 via direct DOM inspection: qwen3.8-flash is listed at $0.15/M input and $0.47/M output for the single 0<Token≤1M band and is marked as supporting context caching. QwenCloud's first-party model page (using the same DashScope International endpoint) publishes $0.016/M implicit cached input, $0.20/M explicit cache creation and $0.016/M explicit cache read. Alibaba's context-cache documentation explicitly lists qwen3.8-flash as a Qwen3.8 exception to the generic 10% cache-read rule, so the cached figure is not derived. A full Azure Retail Prices API sweep on 2026-08-28 found no Foundry meter for Qwen3.8-Flash, so this direct-API row is the only lane.
Qwen3.7 Max (Promo)
Current flagship. Effective rate under a limited-time 50% discount (list $2.50/M in, $7.50/M out) covering all four billing items — input, output, explicit cache creation, and explicit cache hits. Discount is officially scheduled to end 2026-08-31; reverts to list ($2.50 in / $0.25 cached / $7.50 out) from 2026-09-01. Single price tier across the full 1M window; thinking and non-thinking modes priced the same.
Direct API
Model Studio (Intl)
1M$2.50CHF 2.01$0.25CHF 0.201$7.50CHF 6.04official
Alibaba Cloud Model Studio pricing page, International endpoint, captured 2026-08-28: list $2.5/$7.5 marked 'Limited-time 50% off'. Cached input derived as 10% of effective input per the official context-cache rule (explicit cache hits). Alibaba Cloud's campaign page ('Qwen3.8-Max is Here') states that the discount runs until August 31, 2026 and applies to input, output, explicit cache creation and explicit cache hit; the campaign was re-verified 2026-08-28.
List price (from September)
List price (from September)$2.50 / $0.25 / $7.50
Qwen3.7 Plus (Promo)
Effective rate under a limited-time 20% discount (list $0.40/$1.60); no promo end date published. Rates shown are ≤256K prompt tokens; 256K–1M bills $0.96/$3.84 effective ($1.20/$4.80 list). Thinking and non-thinking output priced the same.
Direct API
Model Studio (Intl)
1M$0.32CHF 0.258$0.032CHF 0.026$1.28CHF 1.03official
Alibaba Cloud Model Studio pricing page, International endpoint, captured 2026-07-20: list prices marked 'Limited-time 20% off'. Cached input derived as 10% of effective input per the official context-cache rule.
Qwen3.6 Plus
Tiered: rates shown are ≤256K prompt tokens; 256K-1M bills $2/M input, $6/M output.
Direct API
Model Studio (Intl)
1M$0.50CHF 0.403$0.05CHF 0.04$3.00CHF 2.42official
Alibaba Cloud Model Studio pricing page, International endpoint. Cached input derived as 10% of input per the official context-cache doc (explicit hits 10%, creation 125%, implicit hits 20%), which lists this model as supported.
Qwen3.6 Flash
Cheap tier; 50% batch-inference discount also published.
Direct API
Model Studio (Intl)
$0.25CHF 0.201$0.025CHF 0.02$1.50CHF 1.21official
Alibaba Cloud Model Studio pricing page, International endpoint (≤256K tier). Cached input derived as 10% of input per the official context-cache doc, which lists this model as supported.
Qwen3.7 Flash
Full tier ladder: 0-32K $0.030 in / $0.130 out; 32K-256K $0.100 in / $0.400 out (rate shown, matching the band tracked for the Qwen3.6 Flash row); 256K-1M $0.200 in / $0.800 out.
Direct API
Model Studio (Intl)
$0.10CHF 0.081$0.01CHF 0.008$0.40CHF 0.322official
Alibaba Cloud Model Studio pricing page, International endpoint, snapshot qwen3.7-flash-2026-07-15, captured 2026-07-26. Rate shown is the 32K<tokens≤256K band, chosen to match the context band tracked for the existing Qwen3.6 Flash row (≤256K) so the two are directly comparable. Cached input derived as 10% of input per the official context-cache rule (explicit cache hits), same shape as the rest of the Qwen family.
Qwen3.6 Max Preview
Thinking-only preview; rates shown are the ≤128K tier (128K-256K bills $2/$12). Scheduled for deprecation on 2026-10-10; Alibaba lists Qwen3.7 Max as the replacement.
Direct API
Model Studio (Intl)
$1.30CHF 1.05$0.13CHF 0.105$7.80CHF 6.28official
Alibaba Cloud Model Studio pricing page, International endpoint, re-verified 2026-08-01. Cached input derived as 10% of input per the official context-cache doc, which lists this model as supported. Alibaba's official model-lifecycle page schedules qwen3.6-max-preview for deprecation on 2026-10-10 and names qwen3.7-max as the replacement.
Confidenceofficial published pagederived reconciled from billing estimate pattern-inferred
Field notes

Quirks & gotchas

Watch out

No per-token Qwen meter on Foundry

Microsoft Foundry hosts Qwen models (e.g. the Qwen3-VL family) only via Managed Compute: dedicated A100/H100/MI300 GPUs at $4-8 / CHF 3.22-6.44 per compute hour. There is no serverless per-token Qwen listing on the Foundry pricing page — neither native nor Fireworks-hosted — so cost per token depends entirely on your utilization.

Note

Endpoint pricing differs sharply

The Chinese-mainland (Beijing) endpoint is far cheaper than International: Qwen3.6 Plus drops from $0.50 / CHF 0.40 to ~$0.276 / CHF 0.22 input and from $3.00 / CHF 2.42 to ~$1.65 / CHF 1.33 output. The catalog lists International (Singapore) rates as the realistic option for most non-China deployments.

Insight

Qwen3.8 cache rates are model-specific

Model Studio's Context Cache doc states the generic rule — explicit cache hits at 10% of the input rate, cache creation at 125%, and implicit hits at 20% — but explicitly excludes Qwen3.8-Max, Qwen3.8-Flash and Qwen3.8-27B from that shortcut. Their model pages publish exact cache prices, so the catalog uses those official figures; the older Qwen3.7 and Qwen3.6 rows continue to use the derived 10% explicit-cache-read convention.

Watch out

Flagship pricing is promotional

Qwen3.7 Max ($1.25 / CHF 1.01 input, $3.75 / CHF 3.02 output effective) is billed under a 50% discount that officially ends on 2026-08-31; Qwen3.7 Plus remains 20% off with no published end date. If the promos lapse, Max reverts to $2.50/$7.50 and Plus to $0.40/$1.60. Budget against list price for anything long-lived.

Other providers
KimiDeepSeekGLMOpenAI / Azure OpenAIClaudeGeminiGrokMistralMiniMaxEmbeddingsCompare all →