Skip to content
llm-spend
GitHub
Cost anatomy

Qwen3.7 Max (Promo)

Qwen · Model Studio (Intl)Direct API

Fresh input / cached input / output cost breakdown, active and scheduled rates, provenance, same-model deployment markup, and cost-comparable alternatives for this purchasable lane.

Start with a pattern

Illustrative workloads

Pick a useful starting shape, then fine-tune every input below.

Fine-tune

Workload and pricing scenario

Adjust the token shape, cache behavior, time window, and service tier. Every number below updates immediately.

Interactive

Workload cost calculator

60M
210K
90%

Share of input tokens served from cache. Applied only to models with a cache meter.

Time

Rows with a rate that changes by time of day, promo window or service tier are priced for this scenario instead of their flat listed rate; a brand-colored label under the model name names which one applies. Context size for band-priced rows uses the input-token count below.

Picking a specific hour also previews time-of-day rates that are scheduled but have not started billing yet; those rows are labelled · preview with their start instant. Now only ever shows rates that are billable at this moment.

Rate state

What this lane charges right now

Priced under the workload and scenario above. The full published schedule below always shows every variant regardless of which scenario is selected.

List price (from September) — priced under the selected scenario, differs from the flat listed rate

Input / 1M
$2.50
CHF 2.01
Cached input / 1M
$0.25
CHF 0.201
Output / 1M
$7.50
CHF 6.04
Full published schedule
List price (from September)
List price (from September)$2.50 / $0.25 / $7.50
Cost anatomy

Where this workload's cost goes

60M input tokens (90% cache hit) and 210K output tokens, at the resolved rate above.

$15.00
Fresh input · CHF 12.08
$13.50
Cached input · CHF 10.87
$1.57
Output · CHF 1.27
$30.07
Workload total · CHF 24.21

$15.00 + $13.50 + $1.57 = $30.07

Provenance

How much to trust these numbers

InputofficialCached inputderivedOutputofficial

Current flagship. Effective rate under a limited-time 50% discount (list $2.50/M in, $7.50/M out) covering all four billing items — input, output, explicit cache creation, and explicit cache hits. Discount is officially scheduled to end 2026-08-31; reverts to list ($2.50 in / $0.25 cached / $7.50 out) from 2026-09-01. Single price tier across the full 1M window; thinking and non-thinking modes priced the same.

Alibaba Cloud Model Studio pricing page, International endpoint, captured 2026-08-28: list $2.5/$7.5 marked 'Limited-time 50% off'. Cached input derived as 10% of effective input per the official context-cache rule (explicit cache hits). Alibaba Cloud's campaign page ('Qwen3.8-Max is Here') states that the discount runs until August 31, 2026 and applies to input, output, explicit cache creation and explicit cache hit; the campaign was re-verified 2026-08-28.

Direct vs Foundry

Same-model deployment markup

How much more Microsoft Foundry charges for this exact model, over this Direct rate.

Not available — no Microsoft Foundry listing exists for this exact model in the catalog.

Cost-comparable

Other lanes within ±25% of this workload's cost

Same provider first, then ranked by cost proximity, for 60M in / 210K out. This is a cost ranking only — never a quality recommendation.

  • Grok 4.5
    Grok · xAI direct API
    Direct API
    $29.46
    CHF 23.72
    -$0.615 (-2%) vs this lane
  • GLM 5.1
    GLM · Fireworks-hosted
    Foundry ·Data Zone
    $25.70
    CHF 20.69
    -$4.37 (-15%) vs this lane
  • GPT-5.6 Terra
    OpenAI / Azure OpenAI
    Foundry ·Global
    $25.32
    CHF 20.38
    -$4.75 (-16%) vs this lane
  • GPT-5.2 / Codex
    OpenAI / Azure OpenAI
    Foundry ·Data Zone
    $25.18
    CHF 20.27
    -$4.90 (-16%) vs this lane
  • Claude Sonnet 5
    Claude
    Direct API
    $24.90
    CHF 20.04
    -$5.17 (-17%) vs this lane
  • Claude Sonnet 5
    Claude
    Foundry ·Global
    $24.90
    CHF 20.04
    -$5.17 (-17%) vs this lane