Capital & Compute
· ai· china· pricing· deepseek· open-source

China AI Pricing 2026: DeepSeek, Qwen, Kimi API Costs

Current 2026 API and subscription prices for DeepSeek, Qwen, Kimi, GLM, and MiniMax, compared per million tokens and against US models.

By Capital & Compute

Chinese AI models list API prices well below their US counterparts, and the gap depends entirely on which two models you put side by side. DeepSeek V4 Flash starts at $0.14 per million input tokens. At the other end, Kimi K3 lists output at $15. Line up models of the same measured intelligence and the honest gap runs from about 1.6x to 29x on output tokens, not the flat “90 percent cheaper” you see in headlines. Here is what every major Chinese lab charges in 2026, per token and per month, checked against each provider’s own price page.

$0.14
DeepSeek V4 Flash input
per million tokens, the cheapest flagship-class rate here
29x
DeepSeek V4 vs Claude Opus 4.8
same intelligence tier, output price gap
5 labs
cut token prices 50 to 99 percent
across 2026, per press reports

What Chinese AI models cost per token

These are current list rates per million tokens, verified on July 18, 2026 against each provider’s own pricing page. Input is what you pay to send tokens, output is what the model generates, and the cache column is the discounted rate for input the provider has already seen (the number that decides agent economics, because agents resend the same context every turn):

Model (lab) Input Output Cache read
DeepSeek V4 Flash $0.14 $0.28 $0.0028
MiniMax M3 $0.30 $1.20 $0.06
DeepSeek V4 Pro $0.435 $0.87 $0.0036
Kimi K2.7 Code (Moonshot) $0.95 $4.00 $0.19
GLM-5.2 (Zhipu) $1.40 $4.40 $0.26
Qwen3.7 Max (Alibaba) $2.50 $7.50 $0.25
Kimi K3 (Moonshot) $3.00 $15.00 $0.30

Two things get lost when people quote a single “China is X percent cheaper” number.

First, the cache rates. DeepSeek bills a cache hit at $0.0028 per million tokens, roughly a 98 percent discount on its own input rate, which makes a long system prompt or a repeated codebase almost free to resend. For a coding agent that ships the same 40,000-token context on every turn, that column matters more than the sticker input price. It is the same force pulling the whole market down that we tracked in the collapse in cost per token over time.

Second, DeepSeek runs two tiers that get flattened into one “DeepSeek” row elsewhere. V4 Flash is the budget workhorse at $0.14 input. V4 Pro is the stronger model at $0.435. Both undercut anything comparable from a US lab, but they are not the same product, and quoting the Flash price against a US flagship is the kind of apples-to-oranges move that produces the 90-percent-cheaper headline.

China vs US: the real price gap

The 50x claim is real and misleading at the same time. It is real if you pick the cheapest Chinese model and the most expensive US one. It is misleading as a general statement, because most buying decisions compare models that do roughly the same job.

So match them on capability. Using the intelligence composite from our AI model value leaderboard, here is each Chinese model against the US model closest to it in measured smarts, by output price:

Capability-matched output price: Chinese models vs US peersDeepSeek V4 $0.87 vs Claude Opus 4.8 $25; MiniMax M3 $1.20 vs GPT-5.3 Codex $14; GLM-5.2 $4.40 vs GPT-5.6 Terra $15; Kimi K3 $15 vs GPT-5.6 Sol $30; Qwen3.7 Max $7.50 vs Gemini 3.1 Pro $12.US peerChinese model$0$10$20$30DeepSeek V4 vs Claude Opus 4.8$25$0.87MiniMax M3 vs GPT-5.3 Codex$14$1.2GLM-5.2 vs GPT-5.6 Terra$15$4.4Qwen3.7 Max vs Gemini 3.1 Pro$12$7.5Kimi K3 vs GPT-5.6 Sol$30$15
Capability-matched output price: Chinese models vs US peers
ItemUS peerChinese model
DeepSeek V4 vs Claude Opus 4.8$25$0.87
MiniMax M3 vs GPT-5.3 Codex$14$1.2
GLM-5.2 vs GPT-5.6 Terra$15$4.4
Qwen3.7 Max vs Gemini 3.1 Pro$12$7.5
Kimi K3 vs GPT-5.6 Sol$30$15
Output list price per million tokens, US dollars, July 2026. Each row pairs a Chinese model with the US model nearest to it on the intelligence composite. The gold dot is the Chinese model, the slate dot its US peer; the line is the price gap.Source: Provider price pages and the value leaderboard, verified July 2026

Read the pairs and the story changes row by row. DeepSeek V4 sits within four points of Claude Opus 4.8 on intelligence and costs 29 times less to generate output. That is the extreme, and it is genuine. But Qwen3.7 Max against Gemini 3.1 Pro, two models a rounding error apart on capability, is a 1.6x gap. GLM-5.2 against a mid-tier GPT is about 3.4x. The rule of thumb: peer-to-peer, Chinese models run three to five times cheaper, and only at the top of the US price stack does the multiple blow out past 20x. Still decisive. A 3x difference on output tokens rewrites a product’s unit economics. It just is not 50x across the board.

Plot the whole field and the pattern is a cloud that leans one way:

The value frontier: intelligence vs output priceA scatter of intelligence index against output price per million tokens on a log scale. Chinese models (DeepSeek V4, GLM-5.2, Kimi K3, Qwen3.7 Max, MiniMax M3, Kimi K2.7, Qwen3.7 Plus) sit far below US flagships (Claude Opus 4.8, Claude Fable 5, GPT-5.6 Sol, GPT-5.6 Terra, Gemini 3.1 Pro, Grok 4.5) at comparable intelligence.$1$2$3$4$5$6$7$8$9$10$20$30$40$504045505560intelligence index (Price Per Token composite)output price per million tokens, USD (log)Kimi K3Qwen3.7 MaxGLM-5.2Kimi K2.7 CodeQwen3.7 PlusMiniMax M3DeepSeek V4Claude Fable 5GPT-5.6 SolClaude Opus 4.8GPT-5.6 TerraGemini 3.1 ProGrok 4.5
The value frontier: intelligence vs output price
Itemintelligence index (Price Per Token composite)output price per million tokens, USD (log)
DeepSeek V452$0.87
MiniMax M344.4$1.2
Qwen3.7 Plus39$1.28
Kimi K2.7 Code41.9$4
GLM-5.251.1$4.4
Qwen3.7 Max46$7.5
Kimi K357$15
Grok 4.554$6
Gemini 3.1 Pro46.5$12
GPT-5.6 Terra55$15
Claude Opus 4.855.7$25
GPT-5.6 Sol59$30
Claude Fable 559.9$50
Intelligence composite against output price (log scale), July 2026. Chinese models in gold cluster in the low-price band at every capability level; US flagships in slate sit above them.Source: AI model value leaderboard, benchmark composite from Price Per Token, verified July 2026

The three highest-intelligence models are still American. Everything below them, at any given capability level, has a Chinese option sitting at a fraction of the price. Why that cloud looks the way it does, the mixture-of-experts efficiency, the cheap hub-zone power and the open-weight strategy, is its own story: we walked through it in why Chinese AI models are so cheap. This piece is about what you actually pay.

Subscription and coding-plan pricing

Per-token billing is not the only way to buy these models. Every major Chinese lab now sells a flat-rate coding subscription, the same model as Claude Code or a Cursor plan: pay a fixed monthly fee, get a rolling usage window of prompts inside your editor, stop watching the meter. For steady daily use they undercut token billing by a wide margin, which is why they matter to the “china ai pricing” question as much as the API rates do.

Coding plan (lab) Entry price Runs
MiniMax Starter ~$10/mo MiniMax M2.7
Kimi Code (Moderato) $15/mo Kimi K2.6
GLM Coding Plan (Lite) $18/mo GLM-5.2, GLM-5-Turbo, 4.7
Qwen Code (Pro) $50/mo Qwen3.7-Plus, Kimi, others

Two caveats before you pick one. Alibaba retired its cheaper Qwen Lite tier in March 2026, so $50 a month is now the floor for the Qwen plan, and interestingly that Pro plan bundles rival models (Kimi K2.5, GLM-5, MiniMax) alongside Qwen’s own. GLM’s $18 Lite tier gives you roughly 80 prompts per five-hour window, which is plenty for light use and cramped for a full workday. These plans change often, so we track the live cross-vendor comparison, including the US plans, on the AI coding plan pricing pillar.

The 2026 China AI price war, in dates

None of these numbers hold still. The through-line of 2026 has been Chinese labs cutting, then cutting again, and US labs answering:

  1. April 2026

    DeepSeek V4 lands 97 percent below GPT-5.5

    DeepSeek prices its V4 model at rock-bottom rates, reported by the South China Morning Post as 97 percent below OpenAI’s comparable flagship, and pairs it with Huawei Ascend integration.

  2. May 2026

    DeepSeek makes a 75 percent V4-Pro cut permanent

    What started as a promotion becomes the standing rate. Cached input drops toward a few cents per million tokens.

  3. May 2026

    DeepSeek reaches 17 percent of Vercel token volume

    Rest of World reports DeepSeek at 17 percent of tokens on Vercel’s AI platform while collecting about 1 percent of the revenue: adoption running far ahead of monetization.

  4. Across 2026

    Five labs cut token prices 50 to 99 percent

    ByteDance, Tencent, MiniMax, Alibaba and Xiaomi all cut, per press tallies, with Xiaomi reported slashing some API rates by 99 percent.

  5. July 2026

    US labs answer in the mid-tier

    OpenAI positions GPT-5.6 Terra at near-flagship capability for half the top rate; the compression lands exactly where Chinese models compete.

The direction is one-way for now. When a competitor sells adequate capability at a fifth of your price, you either match on the tiers where you overlap or you cede them. That is the squeeze we unpacked in Qwen 3.7 Max vs Claude for coding, where the premium model still wins on the hardest work while the challenger wins on cost per task.

How to actually buy Chinese AI models

Price is the easy part. Access is where the decision actually lives, and you have three routes.

First-party API. Call DeepSeek, Alibaba, Moonshot, Zhipu or MiniMax directly and you get the list prices above. It is the cheapest path and the one most exposed on compliance: your prompts route through infrastructure inside China and its legal regime, which plenty of Western companies exclude by policy.

A Western reseller or inference host. The same open weights are served by US and EU hosts, so you can run DeepSeek or Qwen without sending data to China. You pay a markup over the first-party rate for that, and the AI inference provider directory lists who carries which model at what price.

Self-hosting the open weights. This is the lever that has no equivalent on the US side. DeepSeek ships under an MIT-style license, and Qwen, Kimi, GLM and MiniMax all release open weights of their own (Kimi K3’s open weights are due July 27, 2026). Download them, run them on your own or rented GPUs, and the per-token API price stops applying entirely: your cost becomes hardware and electricity. For teams with steady, high volume, that is often cheaper than any API rate, Chinese or American. It is also why the Hugging Face ecosystem now leans Chinese, and why the whole category is hard to price with a single number.

If you are weighing the other side of the ledger, the models these prices are pressuring, we keep a running view of the best frontier models excluding Chinese labs and the national picture behind the pricing in the AI usage by country data.

Bottom line

For the “china ai pricing” question, the honest answer is a range, not a slogan. Peer-to-peer, Chinese models cost roughly three to five times less per output token than the US model doing the same job; at the extreme, DeepSeek V4 undercuts a same-tier US flagship by nearly 30x. Coding subscriptions start around $10 to $18 a month against $20 for the cheapest US plan. And for steady volume, the open weights take the API price to zero and leave only your compute bill. The catch is compliance and the durability of subsidized rates, both of which are judgment calls no price table settles for you.

Frequently asked questions

How much does the DeepSeek API cost?
As of July 2026, DeepSeek V4 Flash lists at $0.14 per million input tokens and $0.28 output, and V4 Pro at $0.435 input and $0.87 output. Cache hits are billed near $0.003 per million tokens, a roughly 98 percent discount that makes repeated context almost free.
What is the cheapest Chinese AI model?
On list API rates, DeepSeek V4 Flash is the cheapest flagship-class option at $0.14 input and $0.28 output per million tokens. MiniMax M3 is next at $0.30 and $1.20. Smaller open-weight tiers from Qwen and GLM go lower still, and self-hosting any open-weight model removes the per-token price entirely.
Are Chinese AI models really cheaper than ChatGPT and Claude?
Yes, but the gap depends on the match. Compared model-for-model on capability, Chinese models run about three to five times cheaper per output token. Compared cheapest-to-most-expensive, the gap reaches nearly 30x. The flat 90-percent-cheaper headline only holds at the extremes.
Do Chinese AI models have subscription plans?
Yes. Every major lab sells a flat-rate coding plan: MiniMax from about $10 a month, Kimi Code from $15, GLM from $18, and Qwen at $50 for its Pro tier. Each gives a rolling window of prompts inside your editor instead of per-token billing.
Is it safe to use Chinese AI models?
Calling a China-hosted API routes your data through Chinese infrastructure and legal jurisdiction, which many companies exclude by policy. Running the open weights yourself, or through a US or EU inference host, keeps data under your own jurisdiction and is how most Western production deployments use these models.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks