Capital & Compute

AI Price Hikes 2026: The 10 Nobody Warned You About

Model prices fell up to 80 percent in 2026. These 10 went the other way, from a 571 percent DDR5 jump to a cache rate up 67 percent.

· Updated September 12, 2026· ai· pricing· economics· hardware· By Capital & Compute

The cheapest tracked 32GB DDR5 kit cost $59 in April 2025. On 15 August 2026 the same kit read $396, a rise of 571 percent. That is not a forecast, a spot index or an analyst note. It is one retail series on one multi-retailer tracker, read on a single date, and it is the largest consumer price increase in computing this year.

It is also not the one that will cost most people the most money. That is the finding this list is ordered around.

# What went up The move When
10 GitHub Copilot billing Flat fee gains a meter 1 June 2026
9 ChatGPT Business Premium $100 advertised, $120 real 10 August 2026
8 Grok 4.6 cached input +67%, plus a 200K cliff August 2026
7 Gemini 3.8 Flash Every line doubles 1 January 2027
6 DeepSeek V4 Pro cached input +1,114% at peak hours 16 August 2026
5 Nvidia AI servers More than 15% 22 August 2026
4 Apple Macs, iPads, iPhones iPhone 18 Pro to $1,249-$1,299 From 25 June
3 Hetzner and OVHcloud Up to +192% and +652% Through 2026
2 Nvidia RTX street prices +32% to +50% versus January Late July 2026
1 32GB DDR5 memory +571% To August 2026

Why this list exists

The dominant story about AI pricing is collapse. Token prices have fallen roughly 10x per year since 2021, and on 30 July 2026 OpenAI cut GPT-5.6 Luna by 80 percent. Both of those are true.

They are also the half of the picture that gets written about. The other half is that the physical layer underneath every one of those cheap tokens got repriced, and so did several of the rate-card lines that an agent actually leans on. Cheaper models running on more expensive everything is the real shape of 2026.

Start with the rate cards themselves, because that is where the collapse narrative is strongest and therefore where the counter-examples matter most.

2026 per-token rate changes, cuts against risesFive published per-token rate changes in 2026. GPT-5.6 Luna fell 80 percent from $1.00 and $6.00 to $0.20 and $1.20 per million tokens on July 30. GPT-5.6 Terra fell 20 percent. Qwen3.8 Max list pricing fell 20 percent. Grok 4.6 cached input rose 67 percent from $0.30 to $0.50 per million tokens below 200K context. Gemini 3.8 Flash rises 100 percent on every line on January 1, 2027.Price cutPrice riseGPT-5.6 Luna$1.00/$6.00 to $0.20/$1.20, 30 July-80%GPT-5.6 Terra$2.50/$15 to $2.00/$12, 30 July-20%Qwen3.8 Max, list$2.50/$7.50 to $2.00/$6.00-20%Grok 4.6, cached input$0.30 to $0.50 per Mtok, under 200K+67%Gemini 3.8 FlashEvery line doubles, 1 January 2027+100%-100%-80%-60%-40%-20%0%+20%+40%+60%+80%+100%Each bar is a published list-rate change, not a blended or effective rate
2026 per-token rate changes, cuts against rises
StudyMeasured effectMetric and sample
GPT-5.6 Luna-80%$1.00/$6.00 to $0.20/$1.20, 30 July
GPT-5.6 Terra-20%$2.50/$15 to $2.00/$12, 30 July
Qwen3.8 Max, list-20%$2.50/$7.50 to $2.00/$6.00
Grok 4.6, cached input+67%$0.30 to $0.50 per Mtok, under 200K
Gemini 3.8 Flash+100%Every line doubles, 1 January 2027
Published per-token rate changes in 2026, in both directions. The cuts are real and large. So are the rises, and they land on lines that headline comparisons rarely quote.Source: Provider pricing pages, verified 10 September 2026

One rise is missing from that chart because it would flatten every other bar. It is number 6 below.

10. GitHub Copilot put a meter on a flat fee

GitHub Copilot still costs $10 a month. Since 1 June 2026 that $10 is a floor rather than a price. The plan buys a fixed allowance of GitHub AI Credits, and once the allowance is gone, usage bills by the token at a cent per credit on top, under the move to usage-based billing.

No headline number changed, which is exactly why this one is easy to miss. For three years Copilot had the simplest pricing in AI coding. The full arithmetic, including the flex allotment that makes Pro’s $10 buy $15 of credits, is in what Copilot AI Credits really cost.

9. ChatGPT Business Premium advertises $100 and charges $120

OpenAI removed the rolling five-hour usage cap on ChatGPT Business on 10 August 2026, and put the fix behind a new Premium seat at $125 a month billed monthly or $100 billed annually.

The detail that did not make the announcement: ChatGPT Business has carried a two-seat minimum since launch, and that rule did not change. A workspace cannot buy one Premium seat and stop. The real floor to unlock one uncapped seat is one Standard plus one Premium, which is $120 a month billed annually, not $100. The seat math is worked through here.

8. Grok 4.6 raised the line an agent leans on hardest

xAI kept the Grok 4.6 headline at $2 input and $6 output per million tokens. It raised cached input by 67 percent, from $0.30 to $0.50 below 200K context and from $0.60 to $1.00 above it.

Then there is the cliff. Past 200,000 prompt tokens, xAI rebills the entire request at the higher tier, including the tokens below the line. Crossing the threshold does not price the marginal token differently, it doubles the price of every token in the call. For a long-running agent that accumulates context, that is the whole bill. See the 200K price cliff.

7. Gemini 3.8 Flash has an expiry date on the sticker

Gemini 3.8 Flash costs $0.75 per million input tokens, $3.75 output and $0.075 cached. Google’s own Gemini API pricing page marks all three as introductory rates running through 31 December 2026, rising on 1 January 2027 to $1.50, $7.50 and $0.15.

Everything doubles, on one date, on every line. Budget next year on the number in today’s launch coverage and the figure will be wrong by two thirds. Worse, Gemini 3.6 Flash and 3.7 Flash carried the same $0.75 rate and the same expiry, so the pattern is now three models deep. The crossover math is here.

6. DeepSeek raised cached input 1,114 percent in three days

DeepSeek published its agent harness under an MIT licence on 13 August 2026, free. Three days later it repriced the API that harness runs on, and the line it raised hardest is the one a harness leans on hardest. Cached input on deepseek-v4-pro went from $0.003625 to $0.044 per million tokens at peak hours, a rise of 1,114 percent.

Two changes landed at once, which is why the increase exceeds the pre-announced “2x at peak”. The base rate itself moved roughly 2.3x, and peak hours then double that base. DeepSeek had pre-announced only the second half. Under the old flat card, cached reads were 12.2 percent of a twenty-turn session. Under the peak card they are 31.2 percent. Full teardown in the DeepSeek harness post.

5. Nvidia told its biggest customers AI servers cost more than 15 percent extra

Nvidia notified its largest customers that prices of servers containing its AI chips are rising more than 15 percent in many cases, with increases taking effect on systems shipped early next year and covering the flagship Vera Rubin and Grace Blackwell generations. It was first reported by Bloomberg on 22 August 2026 and confirmed in Fortune’s reporting citing people familiar with the process. The stated cause is soaring memory chip costs.

This is the only item on the list most readers will never see on an invoice, and the one most likely to reach them anyway. Microsoft, Google, Oracle, Amazon and Meta were among those notified. Inference capacity that costs more to build does not stay cheap to rent indefinitely.

4. Apple ended the exemption that mattered most

Apple raised Mac and iPad prices on 25 June 2026, as Reuters reported, citing memory costs. For anyone running models locally, that is a direct tax on the only spec that matters, which is covered in what the unified-memory tax means.

Phones were the last holdout, and that is ending too. The iPhone 18 Pro is expected to start between $1,249 and $1,299 against $1,099 for the iPhone 17 Pro, per TrendForce estimates summarized by iDropNews on 4 September 2026. Memory costs for a 256GB Pro model in the third quarter are expected to land nearly 400 percent higher than a year earlier. The BOM math is here.

3. Your hosting bill was repriced twice, and the average hid it

Hetzner raised prices three separate times in 2026 and published a percentage for none of them. Computing the change from its own before-and-after table for the 15 June adjustment gives the finding: Arm-based CAX plans rose about 31 to 33 percent while x86 CPX and CCX plans rose between 94 and 192 percent, on the same day, in the same data centres. That spread is the memory shortage, isolated. See all three waves.

OVHcloud raised prices twice, and the widely quoted “9 to 11 percent” describes only the first. In that same March announcement the VPS 2026 range moved about 44 to 48 percent. The August wave graded by hardware generation, and memory sold as an option rose 68 to 652 percent on new orders. Founder Octave Klaba said OVHcloud paid six times more for RAM in June 2026 than in June 2025. Both waves are broken out in the OVHcloud post, and the general mechanism in why your VPS price is going up.

2. Graphics cards absorbed three kit hikes in seven months

Nvidia passed memory costs to board partners three times in 2026. TrendForce reported January at 10 to 15 percent, with 16GB-and-above models at 15 to 20 percent. May was narrow and hit the RTX 5090 line only. Late July was the broad one at 20 to 30 percent, reaching every GeForce RTX model.

Those compound. A mid-range card that missed the May increase is roughly 32 to 50 percent more expensive than in January, before a retailer adds anything. An Asus TUF RTX 5080 went from $1,199 to $1,595, or 33 percent, which sits at the bottom of that band. The unabsorbed remainder is the risk to anyone waiting. Full chain in GPU prices in 2026, with current levels on the GPU price tracker.

1. Memory, by a margin nothing else approaches

The cheapest tracked 2x16GB DDR5-6000 kit sat at $59 in April 2025 and read $396 on 15 August 2026, per WhereIsMyRAM’s multi-retailer price chart. That is 571 percent. The tracker resamples its own history between reads, so the figures here are that dated read; a later read on 11 September 2026 put the same kit at $425.

The shape matters as much as the size. On the September 2026 read, the cheap tier held near $195 through January and February 2026, then stepped to roughly $380 in early March. There was a window and it shut inside a quarter. Storage followed the same logic: mainstream NAND contract prices rose 33 to 38 percent in the first quarter of 2026 and 70 to 75 percent in the second, because enterprise SSDs took 48 percent of all NAND bits shipped against 26 percent a year earlier.

Background on the cause is in why RAM is so expensive and why SSDs are so expensive, with live levels on the memory price tracker and the SSD tracker.

The percentage is not the bill

Here is why the ordering above is not the ordering that matters. Take the two best-documented increases on the list, each measured as one continuous series, and apply them to a single upgrade.

What two 2026 price increases add to one upgradeA waterfall chart bridging the cost of a 32GB DDR5 kit plus an Asus TUF RTX 5080 from $1,258 in 2025 to $1,991 in 2026. The DDR5 kit rose 571 percent, adding $337. The graphics card rose 33 percent, adding $396. The pair is 58 percent more expensive, and the smaller percentage contributes the larger dollar amount.$0$500$1,000$1,500$2,000The pair, 2025DDR5 kit $59, RTX 5080 $1,199$1,258DDR5 kit, plus 571%April 2025 to 15 August 2026+$337RTX 5080, plus 33%Winter 2025-26 to July 2026+$396The pair, 2026A 58 percent rise on the two parts$1,991
What two 2026 price increases add to one upgrade
StepChangeRunning total
The pair, 2025 (DDR5 kit $59, RTX 5080 $1,199)$1,258$1,258
DDR5 kit, plus 571% (April 2025 to 15 August 2026)+$337$1,595
RTX 5080, plus 33% (Winter 2025-26 to July 2026)+$396$1,991
The pair, 2026 (A 58 percent rise on the two parts)$1,991$1,991
The 571 percent part adds less money than the 33 percent part. Both figures are single continuous series, not blended baskets.Source: WhereIsMyRAM DDR5 series and reported RTX 5080 street prices

The memory increase is seventeen times larger in percentage terms and contributes $59 less in cash. Percentages describe what happened to a component. Dollars describe what happens to a buyer, and the two rank differently whenever the starting prices differ by an order of magnitude. Any list that ranks 2026 by percentage, including the countdown above, is ranking the wrong column for a purchasing decision.

+571%
Largest consumer component rise
32GB DDR5 kit, April 2025 to August 2026
+1,114%
Largest per-token rise
DeepSeek V4 Pro cached input, at peak
-80%
Largest per-token cut
GPT-5.6 Luna, 30 July 2026

The hike that got cancelled

One scheduled increase did not happen, and leaving it out would make this list dishonest.

Claude Sonnet 5 launched at $2 and $10 per million tokens, described as introductory pricing through 31 August 2026, with a rise to $3 and $15 scheduled for 1 September. Anthropic cancelled it. Its pricing documentation now states that the introductory rate “is now the standard price” and that “the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.”

That is a 50 percent increase withdrawn on a widely used model, and it is the strongest single piece of evidence that competitive pressure on per-token rates is still working, even while everything the tokens run on gets more expensive.

The price cut that raised the bill

The last item is the one that explains why per-token comparisons keep misleading people.

Alibaba cut the Qwen3.8 Max list rate 20 percent, from $2.50 and $7.50 to $2.00 and $6.00. Artificial Analysis then ran the same nine-evaluation suite on it and its predecessor. Qwen3.7 Max cost $1,063.86 to evaluate. Qwen3.8 Max cost $1,741.41. The bill rose 64 percent while the unit price fell, which is only possible if the model emits far more tokens to finish the same work. Detail in the Qwen3.8 Max benchmarks.

A cheaper token that takes more tokens is not a cheaper task. That is the same trap as ranking a build by percentage, arriving from the opposite direction, and it is why the only number worth budgeting on is cost per finished unit of work. Current rates for every tracked model sit on the AI model price tracker, and hosting moves on the hosting price tracker.

Common questions

Frequently asked questions

Did AI prices go up or down in 2026?
Both, on different lines. Published per-token rates fell, in some cases sharply: OpenAI cut GPT-5.6 Luna by 80 percent on July 30, 2026, and Anthropic cancelled a scheduled 50 percent rise on Claude Sonnet 5. Meanwhile the physical layer rose: 32GB DDR5 kits are up 571 percent, Nvidia AI server prices more than 15 percent, and some hosting plans up to 192 percent.
What was the biggest price increase in AI in 2026?
By percentage, DeepSeek cached input on deepseek-v4-pro, which rose 1,114 percent at peak hours on August 16, 2026, from $0.003625 to $0.044 per million tokens. The largest consumer component increase is the cheapest tracked 32GB DDR5 kit at 571 percent, from $59 in April 2025 to $396 in August 2026.
Why is my AI bill going up when model prices are falling?
Three reasons. Headline input and output rates are not the only lines on a rate card, and cached input and long-context tiers rose while headlines held. Introductory rates expire, as Gemini 3.8 Flash does on January 1, 2027. And a cheaper token is not a cheaper task: Qwen3.8 Max cut its list price 20 percent while the cost of running the same evaluation suite rose 64 percent.
Is the memory shortage causing all of this?
Most of it. Memory is the stated cause of the Nvidia server increase, the Apple hardware increases, the Hetzner and OVHcloud hosting increases, and the GPU kit increases, because it is the dominant line in the bill of materials for each. OVHcloud said it paid six times more for RAM in June 2026 than in June 2025. The per-token rises have separate causes.
Should I buy hardware now or wait for prices to fall?
The forward view is genuinely unknown past this quarter. TrendForce had published no quantitative fourth-quarter 2026 DRAM or NAND forecast as of September 8, 2026. The leading indicator has softened, with 512Gb TLC NAND wafer spot about 8 percent below its March peak, and spot leads contract by a quarter or two. Storage is the likeliest first category to stop rising.

Sources

  • GitHub (2026). GitHub Copilot is moving to usage-based billing. GitHub Blog. Verified 10 September 2026. github.blog
  • Google (2026). Gemini API pricing. Official documentation; introductory rates through 31 December 2026 and standard rates from 1 January 2027. Verified 10 September 2026. ai.google.dev
  • Anthropic (2026). Pricing. Official documentation; the note cancelling the 1 September 2026 Claude Sonnet 5 increase. Verified 10 September 2026. platform.claude.com
  • Bloomberg (2026). Nvidia Customers Notified About AI-Related Price Hikes Above 15% (22 August 2026). Secondary, original reporting. Verified 10 September 2026. bloomberg.com
  • Fortune (2026). Nvidia customers notified about AI-related price hikes above 15% (22 August 2026). Secondary, citing people familiar with the process. Verified 10 September 2026. fortune.com
  • Reuters (2026). Apple raises prices on MacBooks, iPads as memory costs skyrocket (25 June 2026). Secondary. Verified 10 September 2026. reuters.com
  • TrendForce News (2026). VRAM Shortage Reportedly Drives NVIDIA’s RTX 50 Price Hikes at Taiwan’s GPU Card Makers This Month (20 January 2026). Secondary, citing Commercial Times. Verified 10 September 2026. trendforce.com
  • WhereIsMyRAM / DropReference (2026). US DDR5 price chart (2x16GB DDR5-6000 CL30 tier). Secondary, multi-retailer aggregated tracker. Verified 10 September 2026. whereismyram.com
  • Counterpoint Research (2026). Server-led eSSDs hit 48 percent of NAND shipments. Secondary, industry analyst. Verified 10 September 2026. counterpointresearch.com
  • iDropNews (2026). iPhone 18 Pro Max price hike estimates (4 September 2026), summarizing TrendForce component-cost estimates. Secondary. Verified 10 September 2026. idropnews.com

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to AI costs