Capital & Compute
· ai· pricing· economics

Cost Per Token Over Time: The AI Price Collapse

How AI cost per token fell roughly 10x per year since 2021, why it collapsed, and why your per-task bill did not drop nearly as fast.

By Capital & Compute

The cost per token of AI has fallen faster than almost any price in computing history. To buy the quality of the March 2023 GPT-4 today, you pay about $0.20 per million output tokens instead of $60, a 300-fold drop in a little over three years. The rate the whole industry quotes is cleaner still: for a model of equivalent performance, the price falls by roughly 10x every year.

That is the headline. The catch, and the reason this post exists, is that the number on your invoice has not fallen 10x a year. Per-token prices collapsed. Per-task bills did not. This is the gap between the two that matters, and it is where most cost planning goes wrong.

How far cost per token has fallen

Two prices tell the story, and they move in opposite ways. One is the price of the frontier flagship: whatever model is called the best that year. The other is the price of a fixed capability: the cheapest model that matches or beats the original GPT-4. Plotting both on a logarithmic axis (each gridline is 10x the one below) shows the split clearly.

AI cost per token over time: frontier flagship vs fixed-capability floorOutput price per million tokens on a log axis. The cheapest GPT-4-class model fell from $60 (GPT-4, March 2023) to about $0.20 (Gemini 2.5 Flash-Lite, 2026), roughly 300x. The frontier flagship line held near $60 (GPT-3 davinci 2021, GPT-4 2023, OpenAI o1 2024) and only recently halved to $30 (GPT-5.6 Sol, 2026).$0.10$1$10$100Nov 2021Mar 2023May 2024Dec 2024Jul 2026Cheapest GPT-4-class model$0.20Frontier flagship$30
AI cost per token over time: frontier flagship vs fixed-capability floor
DateCheapest GPT-4-class modelFrontier flagship
2021-11-01$60
2023-03-01$60$60
2024-05-01$15
2024-07-01$0.60
2024-12-01$60
2026-07-01$0.20$30
Output price per million tokens, log scale. The frontier flagship line barely moves. The cost of buying a fixed quality level collapses.Source: a16z LLMflation (Nov 2024), OpenAI launch pricing, Google Gemini pricing, and this site's model registry

The gap between the two lines is the entire point. When GPT-4 launched, the best model and the cheapest way to get GPT-4 quality were the same thing, at $60 per million output tokens. Three years later the frontier still sits at $30 per million, while GPT-4-level quality is available for cents. The frontier price is sticky. The price of yesterday’s frontier is in free fall.

300x
Fall in GPT-4-class token cost
March 2023 to 2026, output
10x
Per year, equivalent quality
a16z LLMflation, Nov 2024
50x
Per year, median across tasks
Epoch AI, March 2025
1000x
Fall in GPT-3-class token cost
2021 to 2024, per a16z

The single most-cited figure comes from a November 2024 Andreessen Horowitz analysis by Guido Appenzeller, Welcome to LLMflation: the cost to run a GPT-3-class model fell from $60 per million tokens in late 2021 to about $0.06 per million by 2024, a roughly 1,000-fold drop in three years. Appenzeller frames the trend as a 10x fall per year for a model of equivalent performance, with cost at fixed quality halving roughly every two months.

Independent measurement agrees on the direction and pushes the numbers further. Epoch AI, in its March 2025 data insight on LLM inference price trends, found prices falling between 9x and 900x per year depending on the task, with a median of 50x. Restricting the data to after January 2024, the median rose to 200x per year. The price to reach GPT-4’s level on a set of PhD-level science questions fell 40x per year.

How fast per-token prices fall, per year, by estimatea16z estimates a 10x fall per year for an equivalent-quality model. Epoch AI measures a 50x median across tasks, rising to 200x for data since January 2024. The estimates differ because they measure different quality bars, but all describe order-of-magnitude annual declines.0x50x100x150x200xa16z: equivalent-quality model10xEpoch: median across tasks50xEpoch: median since Jan 2024200x
How fast per-token prices fall, per year, by estimate
ItemValue
a16z: equivalent-quality model10x
Epoch: median across tasks50x
Epoch: median since Jan 2024200x
Annual fold-decline in cost per token, by estimate. The three bars measure different things: a16z tracks a single fixed performance level, Epoch aggregates many task-specific quality bars. Both point the same way.Source: a16z LLMflation (Nov 2024) and Epoch AI (March 2025)

Why cost per token collapsed

The fall is not one thing. Four forces stack:

Hardware and serving efficiency. Newer accelerators, better batching, speculative decoding, quantization, and paged attention cut the compute cost of generating a token. This is the slow, compounding floor under everything else.

Distillation and smaller models. A small model trained on a large model’s outputs can reach a capability bar the large model set a year earlier, at a fraction of the parameters and therefore a fraction of the serving cost. GPT-4o mini, launched July 2024 at $0.15 in and $0.60 out per million tokens, was the moment sub-dollar output became normal for a genuinely useful model.

Competition. OpenAI’s own ladder shows the pressure: GPT-4 at $30/$60 in March 2023, GPT-4o at $5/$15 in May 2024, then a further cut months later. Every credible rival that ships forces the incumbents to reprice.

Open weights and Chinese labs. The cheapest capable models on the market are now often open-weight or Chinese-hosted. DeepSeek, Kimi, GLM, and Qwen list frontier-adjacent quality at output prices a fraction of the US flagships. The mechanics of that gap are their own story, covered in why Chinese AI models are so cheap.

Why your bill did not fall 10x a year

Here is the trap. If per-token prices fall 10x a year, it is tempting to assume last year’s workload now costs a tenth as much. It almost never does. Three reasons.

Reasoning models emit far more tokens per task. The 2026 models that top the leaderboards think before they answer, producing long chains of internal reasoning tokens that you pay for. A task that cost 2,000 output tokens on GPT-4 in 2023 can cost 20,000 on a reasoning model today. A 10x drop in price per token is cancelled by a 10x rise in tokens per task. This is the core of the price reversal phenomenon, where the cheaper-per-token model finishes the job at higher total cost.

Context has grown. Bigger context windows invite bigger prompts: whole codebases, long documents, full chat histories re-sent every turn. Input tokens per request have climbed as fast as models allowed them to. Many surprise invoices are input-side, not output-side, as traced in why your AI API bill is so high.

People buy up, not sideways. When quality gets cheaper, most teams do not bank the saving. They move to the newest, most capable (and most expensive) model, because it unlocks work the old one could not do. The frontier price never fell, so buying the frontier never got cheaper.

Cost per token over time: the anchor points

The frontier flagship is the most capable model of its period. The GPT-4-class floor is the cheapest model at or above the original March 2023 GPT-4 quality bar. All prices are list output rates per million tokens at launch.

Date Frontier flagship GPT-4-class floor
Nov 2021 GPT-3 davinci, $60 (GPT-4 not yet released)
March 2023 GPT-4, $60 GPT-4, $60
May 2024 (o1 later in 2024) GPT-4o, $15
July 2024 GPT-4o mini, $0.60
Dec 2024 OpenAI o1, $60
2026 GPT-5.6 Sol, $30 Gemini 2.5 Flash-Lite, $0.20

For current per-token rates across every tracked model, see the AI model release tracker and pricing. For what these rates mean once you turn them into a monthly subscription, see the AI coding plan pricing pillar. And for the case where cheap tokens still lose to owned hardware, see self-hosted LLM cost per token.

What it means for planning

The practical rule: index your budget to capability, not to price per token. If you pin yourself to a fixed quality level, your costs will fall dramatically year over year, and you can plan on it. If you always chase the frontier, expect the top-line price to hold near $30 to $60 per million output tokens indefinitely, because that is where the newest model is always priced.

The deflation is real, and it is the strongest tailwind in the industry. It just does not arrive as a smaller invoice unless you choose to let it.

Frequently asked questions

How much has the cost per token of AI fallen over time?
For a fixed capability level, the cost per token has fallen roughly 10x per year since 2021, according to Andreessen Horowitz. Buying original GPT-4 quality dropped from about $60 to about $0.20 per million output tokens between March 2023 and 2026, roughly 300x. Epoch AI measures a median 50x-per-year fall across tasks, rising to 200x per year for data since January 2024.
Why is my AI bill not falling if prices drop 10x a year?
Because per-token price and per-task cost are different. Reasoning models emit many more tokens per task than older models, and context windows have grown, so tokens-per-request has risen even as price-per-token fell. Most teams also upgrade to the newest, most expensive model rather than bank the saving. The result is a bill that falls far slower than the token price.
Why does the frontier flagship price not fall like everything else?
The newest, most capable model is a scarce good and is priced accordingly. The best model has stayed in the $30 to $60 per million output token range for years: OpenAI o1 launched at the same $60 output price GPT-3 charged at launch. What collapses in price is yesterday-s frontier, once cheaper models match it.
What is LLMflation?
LLMflation is the term Andreessen Horowitz coined in November 2024 for the rapid fall in LLM inference cost: for a model of equivalent performance, cost drops about 10x per year, halving roughly every two months. It describes the deflation in the price of a fixed AI capability, not the price of the frontier.

Sources

  • Appenzeller, Guido (2024). Welcome to LLMflation: LLM inference cost is going down fast. Andreessen Horowitz. a16z.com
  • Epoch AI (2025). LLM inference prices have fallen rapidly but unequally across tasks. Data insight. epoch.ai
  • OpenAI (2024). Hello GPT-4o (GPT-4o pricing $5/$15 per million tokens). openai.com
  • OpenAI (2024). GPT-4o mini: advancing cost-efficient intelligence ($0.15/$0.60 per million tokens). openai.com
  • OpenAI (2023). GPT-4 API general availability ($30/$60 per million tokens at launch). openai.com
  • Google (2026). Gemini API pricing (Gemini 2.5 Flash-Lite $0.05/$0.20 per million tokens). ai.google.dev

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to AI costs