DeepSeek

DeepSeek V4.1 Flash

First CED-architecture sparse MoE, successor to retired V4 Flash

DeepSeek V4.1 Flash is DeepSeek's first sparse-MoE model on the Causal Encoder-Decoder (CED) architecture, served as deepseek-flash, with 1M context, up to 384K output, thinking/non-thinking modes, and vision input. The retired deepseek-v4-flash name still works, and those requests are now served by V4.1 Flash at Flash pricing.

1M Token context1M
Released2026-09
Relays57 sites
Economical API serviceCED architectureMultimodalThinking mode1M context384K output

Relay Comparison

Compare token, per-request, or per-second pricing by relay channel.

MX-AI

mxzzz.xyzToken billing

Price range ¥0.3 - ¥1.2 / 1M tokens

How should DeepSeek V4.1 Flash relay pricing be compared?

This DeepSeek V4.1 Flash pricing page compares official pricing with public prices from 57 listed AI gateways. Token prices are shown in CNY per 1M tokens, while per-request, per-second, and per-character rows use the unit shown in the table. Last updated: 09/20/2026, 23:14.

Data sources
Public price catalogs, official pricing records, and monitoring results.
Metric definitions
Uptime means successful probe response rate, fake-rate signals possible model mismatch or abnormal output risk, and latency is average API response time.
Risk note
Relay gateways are third-party services. Pricing, billing, privacy, and stability can change; start with a small top-up and verify reliability before continued use.