| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| hvoy | Token | ¥0.4 | ¥1.4 | ¥0.115 | ¥0.115 |
Zhipu AI
GLM 5.3 Flash
First native multimodal high-speed coding model in GLM-5
GLM-5.3-Flash is the first native multimodal model in Zhipu's GLM-5 series (320B total / 18B activated), with hybrid linear and sparse attention, image/video/file input, 1M context, and 128K max output. It is available on all GLM Coding Plan tiers at roughly one-third the credit cost of GLM-5.3.
Relay Comparison
Compare token, per-request, or per-second pricing by relay channel.
Price range ¥0.8 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| default | Token | ¥0.8 | ¥2.8 | ¥0.23 | - |
Price range ¥0.48 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| 国产模型 | Token | ¥0.48 | ¥1.68 | ¥0.138 | - |
Price range ¥0.24 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| 国产官方模型 | Token | ¥0.24 | ¥0.84 | ¥0.066 | - |
Price range ¥0.2009 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| 国模2折福利 | Token | ¥0.2009 | ¥0.6698 | ¥0.0402 | ¥0 |
| 国模福利-chat端点 | Token | ¥0.2009 | ¥0.6698 | ¥0.0402 | ¥0 |
| 国模福利-messages端点 | Token | ¥0.2009 | ¥0.6698 | ¥0.0402 | ¥0 |
Price range ¥0.4 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| GLM智谱开源 | Token | ¥0.4 | ¥1.4 | ¥0.115 | - |
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| 国产模型聚合 | Token | ¥0.48 | ¥1.68 | ¥0.138 | - |
Price range ¥0.5 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| glm | Token | ¥0.5 | ¥1.75 | ¥0.1437 | - |
Price range ¥0.144 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| GLM | Token | ¥0.144 | ¥0.504 | ¥0.0414 | - |
Price range ¥0.4 / 1M tokens
| Channel | Billing | Input | Output | Cache hit | Cache write |
|---|---|---|---|---|---|
| 国产模型聚合专区 | Token | ¥0.4 | ¥1.4 | ¥0.115 | - |
How should GLM 5.3 Flash relay pricing be compared?
This GLM 5.3 Flash pricing page compares official pricing with public prices from 58 listed AI gateways. Token prices are shown in CNY per 1M tokens, while per-request, per-second, and per-character rows use the unit shown in the table. Last updated: 09/20/2026, 15:34.
- Data sources
- Public price catalogs, official pricing records, and monitoring results.
- Metric definitions
- Uptime means successful probe response rate, fake-rate signals possible model mismatch or abnormal output risk, and latency is average API response time.
- Risk note
- Relay gateways are third-party services. Pricing, billing, privacy, and stability can change; start with a small top-up and verify reliability before continued use.