China AI Pricing 2026: DeepSeek, Qwen, Kimi API Costs
Current 2026 API and subscription prices for DeepSeek, Qwen, Kimi, GLM, and MiniMax, compared per million tokens and against US models.
By Capital & Compute
Chinese AI models list API prices well below their US counterparts, and the gap depends entirely on which two models you put side by side. DeepSeek V4 Flash starts at $0.14 per million input tokens. At the other end, Kimi K3 lists output at $15. Line up models of the same measured intelligence and the honest gap runs from about 1.6x to 29x on output tokens, not the flat “90 percent cheaper” you see in headlines. Here is what every major Chinese lab charges in 2026, per token and per month, checked against each provider’s own price page.
What Chinese AI models cost per token
These are current list rates per million tokens, verified on July 18, 2026 against each provider’s own pricing page. Input is what you pay to send tokens, output is what the model generates, and the cache column is the discounted rate for input the provider has already seen (the number that decides agent economics, because agents resend the same context every turn):
| Model (lab) | Input | Output | Cache read |
|---|---|---|---|
| DeepSeek V4 Flash | $0.14 | $0.28 | $0.0028 |
| MiniMax M3 | $0.30 | $1.20 | $0.06 |
| DeepSeek V4 Pro | $0.435 | $0.87 | $0.0036 |
| Kimi K2.7 Code (Moonshot) | $0.95 | $4.00 | $0.19 |
| GLM-5.2 (Zhipu) | $1.40 | $4.40 | $0.26 |
| Qwen3.7 Max (Alibaba) | $2.50 | $7.50 | $0.25 |
| Kimi K3 (Moonshot) | $3.00 | $15.00 | $0.30 |
Two things get lost when people quote a single “China is X percent cheaper” number.
First, the cache rates. DeepSeek bills a cache hit at $0.0028 per million tokens, roughly a 98 percent discount on its own input rate, which makes a long system prompt or a repeated codebase almost free to resend. For a coding agent that ships the same 40,000-token context on every turn, that column matters more than the sticker input price. It is the same force pulling the whole market down that we tracked in the collapse in cost per token over time.
Second, DeepSeek runs two tiers that get flattened into one “DeepSeek” row elsewhere. V4 Flash is the budget workhorse at $0.14 input. V4 Pro is the stronger model at $0.435. Both undercut anything comparable from a US lab, but they are not the same product, and quoting the Flash price against a US flagship is the kind of apples-to-oranges move that produces the 90-percent-cheaper headline.
China vs US: the real price gap
The 50x claim is real and misleading at the same time. It is real if you pick the cheapest Chinese model and the most expensive US one. It is misleading as a general statement, because most buying decisions compare models that do roughly the same job.
So match them on capability. Using the intelligence composite from our AI model value leaderboard, here is each Chinese model against the US model closest to it in measured smarts, by output price:
| Item | US peer | Chinese model |
|---|---|---|
| DeepSeek V4 vs Claude Opus 4.8 | $25 | $0.87 |
| MiniMax M3 vs GPT-5.3 Codex | $14 | $1.2 |
| GLM-5.2 vs GPT-5.6 Terra | $15 | $4.4 |
| Qwen3.7 Max vs Gemini 3.1 Pro | $12 | $7.5 |
| Kimi K3 vs GPT-5.6 Sol | $30 | $15 |
Read the pairs and the story changes row by row. DeepSeek V4 sits within four points of Claude Opus 4.8 on intelligence and costs 29 times less to generate output. That is the extreme, and it is genuine. But Qwen3.7 Max against Gemini 3.1 Pro, two models a rounding error apart on capability, is a 1.6x gap. GLM-5.2 against a mid-tier GPT is about 3.4x. The rule of thumb: peer-to-peer, Chinese models run three to five times cheaper, and only at the top of the US price stack does the multiple blow out past 20x. Still decisive. A 3x difference on output tokens rewrites a product’s unit economics. It just is not 50x across the board.
Plot the whole field and the pattern is a cloud that leans one way:
| Item | intelligence index (Price Per Token composite) | output price per million tokens, USD (log) |
|---|---|---|
| DeepSeek V4 | 52 | $0.87 |
| MiniMax M3 | 44.4 | $1.2 |
| Qwen3.7 Plus | 39 | $1.28 |
| Kimi K2.7 Code | 41.9 | $4 |
| GLM-5.2 | 51.1 | $4.4 |
| Qwen3.7 Max | 46 | $7.5 |
| Kimi K3 | 57 | $15 |
| Grok 4.5 | 54 | $6 |
| Gemini 3.1 Pro | 46.5 | $12 |
| GPT-5.6 Terra | 55 | $15 |
| Claude Opus 4.8 | 55.7 | $25 |
| GPT-5.6 Sol | 59 | $30 |
| Claude Fable 5 | 59.9 | $50 |
The three highest-intelligence models are still American. Everything below them, at any given capability level, has a Chinese option sitting at a fraction of the price. Why that cloud looks the way it does, the mixture-of-experts efficiency, the cheap hub-zone power and the open-weight strategy, is its own story: we walked through it in why Chinese AI models are so cheap. This piece is about what you actually pay.
Subscription and coding-plan pricing
Per-token billing is not the only way to buy these models. Every major Chinese lab now sells a flat-rate coding subscription, the same model as Claude Code or a Cursor plan: pay a fixed monthly fee, get a rolling usage window of prompts inside your editor, stop watching the meter. For steady daily use they undercut token billing by a wide margin, which is why they matter to the “china ai pricing” question as much as the API rates do.
| Coding plan (lab) | Entry price | Runs |
|---|---|---|
| MiniMax Starter | ~$10/mo | MiniMax M2.7 |
| Kimi Code (Moderato) | $15/mo | Kimi K2.6 |
| GLM Coding Plan (Lite) | $18/mo | GLM-5.2, GLM-5-Turbo, 4.7 |
| Qwen Code (Pro) | $50/mo | Qwen3.7-Plus, Kimi, others |
Two caveats before you pick one. Alibaba retired its cheaper Qwen Lite tier in March 2026, so $50 a month is now the floor for the Qwen plan, and interestingly that Pro plan bundles rival models (Kimi K2.5, GLM-5, MiniMax) alongside Qwen’s own. GLM’s $18 Lite tier gives you roughly 80 prompts per five-hour window, which is plenty for light use and cramped for a full workday. These plans change often, so we track the live cross-vendor comparison, including the US plans, on the AI coding plan pricing pillar.
The 2026 China AI price war, in dates
None of these numbers hold still. The through-line of 2026 has been Chinese labs cutting, then cutting again, and US labs answering:
April 2026
DeepSeek V4 lands 97 percent below GPT-5.5
DeepSeek prices its V4 model at rock-bottom rates, reported by the South China Morning Post as 97 percent below OpenAI’s comparable flagship, and pairs it with Huawei Ascend integration.
May 2026
DeepSeek makes a 75 percent V4-Pro cut permanent
What started as a promotion becomes the standing rate. Cached input drops toward a few cents per million tokens.
May 2026
DeepSeek reaches 17 percent of Vercel token volume
Rest of World reports DeepSeek at 17 percent of tokens on Vercel’s AI platform while collecting about 1 percent of the revenue: adoption running far ahead of monetization.
Across 2026
Five labs cut token prices 50 to 99 percent
ByteDance, Tencent, MiniMax, Alibaba and Xiaomi all cut, per press tallies, with Xiaomi reported slashing some API rates by 99 percent.
July 2026
US labs answer in the mid-tier
OpenAI positions GPT-5.6 Terra at near-flagship capability for half the top rate; the compression lands exactly where Chinese models compete.
The direction is one-way for now. When a competitor sells adequate capability at a fifth of your price, you either match on the tiers where you overlap or you cede them. That is the squeeze we unpacked in Qwen 3.7 Max vs Claude for coding, where the premium model still wins on the hardest work while the challenger wins on cost per task.
How to actually buy Chinese AI models
Price is the easy part. Access is where the decision actually lives, and you have three routes.
First-party API. Call DeepSeek, Alibaba, Moonshot, Zhipu or MiniMax directly and you get the list prices above. It is the cheapest path and the one most exposed on compliance: your prompts route through infrastructure inside China and its legal regime, which plenty of Western companies exclude by policy.
A Western reseller or inference host. The same open weights are served by US and EU hosts, so you can run DeepSeek or Qwen without sending data to China. You pay a markup over the first-party rate for that, and the AI inference provider directory lists who carries which model at what price.
Self-hosting the open weights. This is the lever that has no equivalent on the US side. DeepSeek ships under an MIT-style license, and Qwen, Kimi, GLM and MiniMax all release open weights of their own (Kimi K3’s open weights are due July 27, 2026). Download them, run them on your own or rented GPUs, and the per-token API price stops applying entirely: your cost becomes hardware and electricity. For teams with steady, high volume, that is often cheaper than any API rate, Chinese or American. It is also why the Hugging Face ecosystem now leans Chinese, and why the whole category is hard to price with a single number.
If you are weighing the other side of the ledger, the models these prices are pressuring, we keep a running view of the best frontier models excluding Chinese labs and the national picture behind the pricing in the AI usage by country data.
Bottom line
For the “china ai pricing” question, the honest answer is a range, not a slogan. Peer-to-peer, Chinese models cost roughly three to five times less per output token than the US model doing the same job; at the extreme, DeepSeek V4 undercuts a same-tier US flagship by nearly 30x. Coding subscriptions start around $10 to $18 a month against $20 for the cheapest US plan. And for steady volume, the open weights take the API price to zero and leave only your compute bill. The catch is compliance and the durability of subsidized rates, both of which are judgment calls no price table settles for you.
Frequently asked questions
- How much does the DeepSeek API cost?
- As of July 2026, DeepSeek V4 Flash lists at $0.14 per million input tokens and $0.28 output, and V4 Pro at $0.435 input and $0.87 output. Cache hits are billed near $0.003 per million tokens, a roughly 98 percent discount that makes repeated context almost free.
- What is the cheapest Chinese AI model?
- On list API rates, DeepSeek V4 Flash is the cheapest flagship-class option at $0.14 input and $0.28 output per million tokens. MiniMax M3 is next at $0.30 and $1.20. Smaller open-weight tiers from Qwen and GLM go lower still, and self-hosting any open-weight model removes the per-token price entirely.
- Are Chinese AI models really cheaper than ChatGPT and Claude?
- Yes, but the gap depends on the match. Compared model-for-model on capability, Chinese models run about three to five times cheaper per output token. Compared cheapest-to-most-expensive, the gap reaches nearly 30x. The flat 90-percent-cheaper headline only holds at the extremes.
- Do Chinese AI models have subscription plans?
- Yes. Every major lab sells a flat-rate coding plan: MiniMax from about $10 a month, Kimi Code from $15, GLM from $18, and Qwen at $50 for its Pro tier. Each gives a rolling window of prompts inside your editor instead of per-token billing.
- Is it safe to use Chinese AI models?
- Calling a China-hosted API routes your data through Chinese infrastructure and legal jurisdiction, which many companies exclude by policy. Running the open weights yourself, or through a US or EU inference host, keeps data under your own jurisdiction and is how most Western production deployments use these models.
Sources
- DeepSeek (2026). Models and Pricing. DeepSeek API documentation. https://api-docs.deepseek.com/quick_start/pricing
- Alibaba Cloud (2026). Model Studio model pricing. Documentation. https://www.alibabacloud.com/help/en/model-studio/model-pricing
- Alibaba Cloud (2026). Model Studio Coding Plan. Documentation. https://www.alibabacloud.com/help/en/model-studio/coding-plan
- Moonshot AI (2026). Kimi platform pricing. Documentation. https://platform.kimi.ai/docs/pricing
- Moonshot AI (2026). Kimi K2.6 pricing. https://www.kimi.com/resources/kimi-k2-6-pricing
- Zhipu / Z.ai (2026). Pricing overview and Coding Plan (DevPack). Documentation. https://docs.z.ai/guides/overview/pricing
- MiniMax (2026). Pay-as-you-go model pricing. Platform documentation. https://platform.minimax.io/docs/guides/pricing-paygo
- Hugging Face (2026). State of Open Source on Hugging Face: Spring 2026. https://huggingface.co/blog/huggingface/state-of-os-hf-spring-2026
- South China Morning Post (2026). China DeepSeek prices new V4 AI model 97 percent below OpenAI GPT-5.5 (as reported). https://www.scmp.com/tech/tech-trends/article/3351595/chinas-deepseek-prices-new-v4-ai-model-97-below-openais-gpt-55
- Rest of World (2026). Low-cost Chinese AI models like DeepSeek gain traction in the U.S. (as reported). https://restofworld.org/2026/when-americans-choose-chinese-ai/
- AI Weekly (2026). Five Chinese AI labs cut token prices up to 99 percent (as reported). https://aiweekly.co/alerts/five-chinese-ai-labs-cut-token-prices-up-to-99