Zhipu AI

GLM 5.3 Flash

First native multimodal high-speed coding model in GLM-5

GLM-5.3-Flash is the first native multimodal model in Zhipu's GLM-5 series (320B total / 18B activated), with hybrid linear and sparse attention, image/video/file input, 1M context, and 128K max output. It is available on all GLM Coding Plan tiers at roughly one-third the credit cost of GLM-5.3.

1M Token context1M
Released2026-08
Relays58 sites
multimodaltool useopen weights1M context128K output

Relay Comparison

Compare token, per-request, or per-second pricing by relay channel.

Price range ¥0.2009 / 1M tokens

How should GLM 5.3 Flash relay pricing be compared?

This GLM 5.3 Flash pricing page compares official pricing with public prices from 58 listed AI gateways. Token prices are shown in CNY per 1M tokens, while per-request, per-second, and per-character rows use the unit shown in the table. Last updated: 09/20/2026, 15:34.

Data sources
Public price catalogs, official pricing records, and monitoring results.
Metric definitions
Uptime means successful probe response rate, fake-rate signals possible model mismatch or abnormal output risk, and latency is average API response time.
Risk note
Relay gateways are third-party services. Pricing, billing, privacy, and stability can change; start with a small top-up and verify reliability before continued use.