Moonshot AI

Kimi K3

Kimi flagship model with 2.8T parameters, native vision, and a 1M context window

Kimi K3 is Moonshot AI's most capable flagship model, with 2.8 trillion parameters, Kimi Delta Attention, Attention Residuals, sparse MoE, native vision, and a 1,048,576-token context window. It targets long-horizon coding, knowledge work, and reasoning, with tool calling, structured output, Partial Mode, and automatic context caching.

1,048,576-token context1M
Released2026-07-17
Relays42 sites
2.8T parametersNative visionLong-horizon codingKnowledge workReasoningTool use1M context

Moonshot AI Official Pricing

CNY
Updated: 2026-07-17T10:50:25.324+08:00Source

Input

¥20/ 1M tokens

Official flat pricing with no context-length tiers: CNY 20 uncached input, CNY 2 cached input, and CNY 100 output per 1M tokens; the context window is 1,048,576 tokens.

Output

¥100/ 1M tokens

Official flat pricing with no context-length tiers: CNY 20 uncached input, CNY 2 cached input, and CNY 100 output per 1M tokens; the context window is 1,048,576 tokens.

Cache read

¥2/ 1M tokens

Official flat pricing with no context-length tiers: CNY 20 uncached input, CNY 2 cached input, and CNY 100 output per 1M tokens; the context window is 1,048,576 tokens.

Official flat pricing with no context-length tiers: CNY 20 uncached input, CNY 2 cached input, and CNY 100 output per 1M tokens; the context window is 1,048,576 tokens.

Relay Comparison

Compare token, per-request, or per-second pricing by relay channel.

How should Kimi K3 relay pricing be compared?

This Kimi K3 pricing page compares official pricing with public prices from 42 listed AI gateways. Token prices are shown in CNY per 1M tokens, while per-request, per-second, and per-character rows use the unit shown in the table. Last updated: 08/06/2026, 23:13.

Data sources
Public price catalogs, official pricing records, and monitoring results.
Metric definitions
Uptime means successful probe response rate, fake-rate signals possible model mismatch or abnormal output risk, and latency is average API response time.
Risk note
Relay gateways are third-party services. Pricing, billing, privacy, and stability can change; start with a small top-up and verify reliability before continued use.