Alibaba Cloud

Qwen3.7 Max

Qwen flagship model for agents, complex coding, and long-horizon reasoning

Qwen3.7 Max is Alibaba Cloud Model Studio's flagship Qwen model. It supports thinking and non-thinking modes, up to 1M input tokens, context caching, and Batch calls, with deployments across mainland China, Singapore, Japan, Germany, and the United States. The current qwen3.7-max alias is documented as equivalent to qwen3.7-max-2026-05-20.

1M input context1M
Released2026-05-20
Relays25 sites
Agentic capabilityComplex codingLong-horizon reasoningThinking modeContext caching1M context

Alibaba Cloud Official Pricing

CNY
Updated: 2026-07-17T10:50:25.324+08:00Source

official-current-promo · Input

¥6/ 1M tokens

Current limited-time 50% mainland China price: CNY 6 input and CNY 18 output per 1M tokens. Explicit cache creation is 125% of input (CNY 7.5) and cache hits are 10% (CNY 0.6). The official page does not publish an end date; recheck before production execution.

official-current-promo · Output

¥18/ 1M tokens

Current limited-time 50% mainland China price: CNY 6 input and CNY 18 output per 1M tokens. Explicit cache creation is 125% of input (CNY 7.5) and cache hits are 10% (CNY 0.6). The official page does not publish an end date; recheck before production execution.

official-current-promo · Cache read

¥0.6/ 1M tokens

Current limited-time 50% mainland China price: CNY 6 input and CNY 18 output per 1M tokens. Explicit cache creation is 125% of input (CNY 7.5) and cache hits are 10% (CNY 0.6). The official page does not publish an end date; recheck before production execution.

official-current-promo · Cache write

¥7.5/ 1M tokens

Current limited-time 50% mainland China price: CNY 6 input and CNY 18 output per 1M tokens. Explicit cache creation is 125% of input (CNY 7.5) and cache hits are 10% (CNY 0.6). The official page does not publish an end date; recheck before production execution.

official · Input

¥12/ 1M tokens

Official mainland China list price: CNY 12 input and CNY 36 output per 1M tokens. Explicit cache creation is 125% of input (CNY 15), explicit cache hits are 10% (CNY 1.2), and implicit cache hits are 20% (CNY 2.4).

official · Output

¥36/ 1M tokens

Official mainland China list price: CNY 12 input and CNY 36 output per 1M tokens. Explicit cache creation is 125% of input (CNY 15), explicit cache hits are 10% (CNY 1.2), and implicit cache hits are 20% (CNY 2.4).

official · Cache read

¥1.2/ 1M tokens

Official mainland China list price: CNY 12 input and CNY 36 output per 1M tokens. Explicit cache creation is 125% of input (CNY 15), explicit cache hits are 10% (CNY 1.2), and implicit cache hits are 20% (CNY 2.4).

official · Cache write

¥15/ 1M tokens

Official mainland China list price: CNY 12 input and CNY 36 output per 1M tokens. Explicit cache creation is 125% of input (CNY 15), explicit cache hits are 10% (CNY 1.2), and implicit cache hits are 20% (CNY 2.4).

Current limited-time 50% mainland China price: CNY 6 input and CNY 18 output per 1M tokens. Explicit cache creation is 125% of input (CNY 7.5) and cache hits are 10% (CNY 0.6). The official page does not publish an end date; recheck before production execution.

Relay Comparison

Compare token, per-request, or per-second pricing by relay channel.

How should Qwen3.7 Max relay pricing be compared?

This Qwen3.7 Max pricing page compares official pricing with public prices from 25 listed AI gateways. Token prices are shown in CNY per 1M tokens, while per-request, per-second, and per-character rows use the unit shown in the table. Last updated: 08/06/2026, 23:13.

Data sources
Public price catalogs, official pricing records, and monitoring results.
Metric definitions
Uptime means successful probe response rate, fake-rate signals possible model mismatch or abnormal output risk, and latency is average API response time.
Risk note
Relay gateways are third-party services. Pricing, billing, privacy, and stability can change; start with a small top-up and verify reliability before continued use.