Capital & Compute

New AI Models Released in September 2026: Prices

Every AI model released in September 2026, with dates, verified per-token prices and a primary source for each. Updated through the month as releases land.

· Updated September 6, 2026· ai· pricing· economics· By Capital & Compute

Seven models have been released in September 2026 so far, across five labs and three days. Anthropic opened on September 1 with Claude Fable 5.1 and its invitation-only twin Claude Mythos 5.1. Google and Meta both shipped on September 2: Gemini 3.8 Flash with a restricted Gemini 3.8 Flash Cyber variant, and Muse Spark 1.3. Alibaba refreshed its flagship the same day with the Qwen3.8-Max-0902 snapshot. OpenAI released GPT-6 Astra on September 3.

For three days the story was that nobody was changing a price. Fable 5.1 holds Fable 5 rates, Gemini 3.8 Flash holds Gemini 3.7 Flash rates, Muse Spark 1.3 holds Muse Spark 1.2 rates, and Qwen3.8-Max-0902 holds the August Qwen3.8-Max rates. What moved in those six launches was a cache rate, a scheduled expiry, and the capability delivered per dollar.

Then OpenAI broke it. GPT-6 Astra lists at $10 input and $50 output per million tokens, two and a half times the $4 and $20 that GPT-5.6 Sol charges, and the first increase at the frontier tier in years. The catch, and the reason this month is more interesting than a price-rise headline, is that Astra is also the cheapest way to finish a task on the one public leaderboard that publishes what a run costs.

  1. September 1, 2026

    Claude Fable 5.1 and Claude Mythos 5.1 released

    Anthropic held Fable 5 rates at $10 input and $50 output per million tokens and cut only the cache read, from $1.00 to $0.25. Mythos 5.1 shares the specifications and the price but stays restricted to invited Project Glasswing participants.

  2. September 1, 2026

    The scheduled Claude Sonnet 5 increase did not happen

    Sonnet 5 launched at $2 and $10 per million tokens, described as introductory pricing through August 31 with a rise to $3 and $15 due on that date. Anthropic cancelled it in August; the pricing documentation now records $2 and $10 as the standard rate.

  3. September 2, 2026

    Google Gemini 3.8 Flash and Gemini 3.8 Flash Cyber released

    Google held the $0.75 and $3.75 per million token rate that Gemini 3.6 Flash and Gemini 3.7 Flash already carry. All three share one expiry: the rate is introductory through December 31, 2026 and doubles on January 1, 2027. The Cyber variant is restricted to the Fairwind Program for governments, critical infrastructure operators and software maintainers.

  4. September 2, 2026

    Meta Muse Spark 1.3 released

    Meta held the $1.25 and $4.25 per million token rate card of Muse Spark 1.2 and moved the independent Artificial Analysis Intelligence Index from 57 to 61 at xhigh effort. A partner-only max configuration scores 62 and stays in limited preview.

  5. September 2, 2026

    Alibaba Qwen3.8-Max-0902 snapshot released

    Qwen held the $2 and $6 per million token rate of the August Qwen3.8-Max and changed only the post-training, aimed at coding and cowork tasks. The qwen3.8-max endpoint upgrades to the snapshot automatically on September 5 with billing unchanged.

  6. September 3, 2026

    OpenAI GPT-6 Astra released

    OpenAI raised the frontier rate for the first time in years, to $10 input and $50 output per million tokens, against $4 and $20 for GPT-5.6 Sol. Prompts over 272,000 tokens are billed at double, with output at $75. It went straight to the top of the independent Terminal-Bench 4.0 board at 58.18%, and costs less per solved task than any model above 20% accuracy.

  7. Still pending

    Google Gemini 3.5 Pro

    Announced and listed as coming soon, with no date and no published price. After GPT-6 Astra shipped it is the only model left in the tracker upcoming table.

Every AI model released in September 2026

Model Provider Released Input / output per Mtok What it is
Claude Fable 5.1 Anthropic Sep 1 $10 / $50 Frontier flagship. Fable 5 rates held, cache read cut to $0.25
Claude Mythos 5.1 Anthropic Sep 1 $10 / $50 Same model, safeguards lifted in some areas. Project Glasswing invitation only
Gemini 3.8 Flash Google Sep 2 $0.75 / $3.75 Workhorse Flash tier. Rate held from 3.7 Flash, and doubles on Jan 1, 2027
Gemini 3.8 Flash Cyber Google Sep 2 Not published Vulnerability-hunting variant. Fairwind Program access only
Muse Spark 1.3 Meta Sep 2 $1.25 / $4.25 Agentic flagship. Rate held from 1.2, index 57 to 61 at xhigh effort
Qwen3.8-Max-0902 Alibaba Sep 2 $2 / $6 Flagship snapshot. Same base and rate, post-training aimed at coding
GPT-6 Astra OpenAI Sep 3 $10 / $50 New frontier flagship. 2.5x the GPT-5.6 Sol rate, and tops Terminal-Bench 4.0

Every rate is read from the provider pricing documentation rather than from an announcement: the Claude Platform pricing documentation, the Gemini API pricing page, the Meta Model API pricing and rate limits page, the QwenCloud model page and the OpenAI API pricing documentation. The two Claude models carry a 1M-token context window with 128K maximum output; Gemini 3.8 Flash carries 1M with 64K output, Muse Spark 1.3 carries 1,048,576 tokens, Qwen3.8-Max-0902 carries 1M with 131K output, and GPT-6 Astra carries 1,050,000 with 128K output.

Claude Fable 5.1: the sticker held, the cache read did not

Fable 5.1 was the month’s first launch and the only one to change a published rate at all. Every headline number matches Fable 5 exactly. Input $10, output $50, cache writes $12.50 for five minutes and $20 for an hour, batch at half price. The single change is the cache read, down from $1.00 to $0.25 per million tokens.

That rate is an exception in Anthropic’s own price list. Every other Claude model charges 10% of its base input rate for a cache hit. Fable 5.1 and Mythos 5.1 charge 2.5%.

Claude cache read rates per million tokens, September 2026A ladder of cache read rates per million tokens across the Claude lineup. Claude Fable 5 charged $1.00 until September 1. Claude Opus 5 charges $0.50. Claude Fable 5.1 charges $0.25, half the Opus 5 rate despite costing twice as much on base input. Claude Sonnet 5 charges $0.20 and Claude Haiku 4.5 charges $0.10.$0.10$0.25$0.50$1.00USD per million cached tokens (log scale)Claude Fable 5 (until Sep 1)$1.00Claude Opus 5$0.50Claude Fable 5.1$0.25Claude Sonnet 5$0.20Claude Haiku 4.5$0.10
Claude cache read rates per million tokens, September 2026
ToolCost per taskMultiple of baseline
Claude Fable 5 (until Sep 1)$1.00-
Claude Opus 5$0.50-
Claude Fable 5.1$0.25-
Claude Sonnet 5$0.20-
Claude Haiku 4.5$0.10-
Cache read rates across the current Claude lineup after September 1. Fable 5.1 undercuts every model above Haiku, including the Opus 5 that costs half as much on base input.Source: Claude Platform pricing documentation

Why that matters to a bill rather than a rate card: a coding agent re-reads the same file tree and conversation history on every step, so most of its input is cache hits. Modeled on the task profiles in the cost-per-task calculator, a long agentic run drops from $9.70 on Fable 5 to $7.68 on Fable 5.1, about 21%, with fresh input and output contributing nothing to the change. A small single-file edit saves only 11%, because less of its input comes from cache.

Anthropic quotes up to about 45% for highly agentic work. That figure needs a cache share above 93%, with essentially no output tokens. The full arithmetic, including why Fable 5.1 still never undercuts Opus 5 on a task that generates output, is in Claude Fable 5.1: pricing, benchmarks, cost.

Gemini 3.8 Flash: the third Flash in a row at the same price

Google’s third Flash release in six weeks lists $0.75 per million input tokens and $3.75 per million output tokens, the same rate Gemini 3.6 Flash and Gemini 3.7 Flash already carry. So does the expiry attached to it. Google’s own pricing page marks all three as introductory through December 31, 2026, rising to $1.50 and $7.50 on January 1, 2027, with cached input going $0.075 to $0.15.

That is the fact worth carrying out of this month. Upgrading generation does not reset the clock, and there is no cheaper Flash generation to retreat to, because all three are the same rate card. A team budgeting 2027 from a September rate card understates a steady workload by two thirds.

Google reports 54.9% on HLE-Verified and says the model beats most larger frontier models on DeepSWE v1.1; both are vendor figures. Artificial Analysis, an independent benchmarking service, scores it 59 on its Intelligence Index at high effort, against 56 for Gemini 3.7 Flash.

Muse Spark 1.3: the rate card that has not moved in three releases

Meta shipped its fourth Muse Spark in five months at $1.25 and $4.25 per million tokens, unchanged since Muse Spark 1.1 in July. The cheaper contributor endpoint, $0.10 and $0.20 per million tokens in exchange for Meta using the prompts and completions sent to it, is unchanged too.

The change is capability per dollar. Artificial Analysis scores the generally available xhigh configuration at 61, up from 57 for Muse Spark 1.2, which puts it level with GPT-5.6 Sol at max effort and ahead of Gemini 3.8 Flash. Meta also reports about 20% fewer tool calls and 25% fewer tokens per coding task than 1.2, a saving that reaches a bill without touching a rate.

Both September 2 launches, the January price cliff, and the modeled twelve-month bills are worked through in Gemini 3.8 Flash and Muse Spark 1.3 pricing.

Qwen3.8-Max-0902: the same flagship, retrained

Alibaba’s September 2 entry is a snapshot, not a new model underneath. Qwen3.8-Max-0902 keeps the August Qwen3.8-Max base, all 2.4 trillion parameters of it, and the $2 and $6 per million token rate, with implicit cache reads at $0.25. What changed is the post-training, extended on coding and on Cowork, the agentic office work Qwen measures with its own CoWorkBench.

The gains on Qwen’s own table are large where the August snapshot was weakest: Terminal-Bench 3.0 from 11.3 to 29.0, JobBench 53.4 to 64.0. Two footnotes travel with that table. Three of its benchmarks are Qwen in-house tests, and the rival column it beats is Fable 5, not the Fable 5.1 that shipped the same week. Treat the direction as signal and the margins as provisional.

The operational detail matters more than the benchmarks. On September 5 the qwen3.8-max endpoint upgrades to this snapshot automatically with billing unchanged, so pinned callers get the new post-training whether they asked for it or not. The snapshot is API-only: no weights were published for it, leaving the August base the most recent downloadable Qwen3.8.

GPT-6 Astra: the first frontier price rise in years

OpenAI ended the no-price-change run on September 3. GPT-6 Astra lists $10 per million input tokens and $50 per million output on the OpenAI API pricing documentation, against $4 and $20 for GPT-5.6 Sol. Cached input is $1 and cache writes are $12.50. Prompts above 272,000 tokens are billed at double for the whole request, with output at $75, and Fast mode is twice the standard rate.

Two things make that less simple than a 2.5x rise. The rate it is measured against is itself temporary: OpenAI flags the Sol figures as promotional through at least November 21, 2026. And $10 and $50 is not a new high. It is exactly what Claude Fable 5.1 charged two days earlier.

The sticker is also the wrong number to judge it by. Terminal-Bench 4.0 publishes the dollar cost of every run alongside the score, and Astra went straight to the top of it at 58.18% across 330 trials. Dividing the published spend by solved trials, the arithmetic used throughout the Terminal-Bench 4.0 analysis here, Astra solves a task for $17.02 at max effort and $11.88 at high effort. Claude Fable 5.1 scores a statistically identical 57.88% and costs $32.69. Claude Opus 5 costs $34.91 for six points less.

The mechanism is token burn, not the rate card. Astra used 1.53 billion tokens across the benchmark. Fable 5.1 used 2.75 billion and Opus 5 used 6.53 billion. A model can charge two and a half times more per token and still be the cheapest way to finish the work.

The release pace

Seven releases in the first three days is a fast start, not a slow month. Set it against the three months before it and the shape is clearer.

Tracked AI model releases by month, June to September 2026A lollipop chart of tracked AI model releases by month. June 2026 had 5 releases, July had 21, August had 12, and September has 7 counted through September 6 with the month still running.0510152025June 20265July 202621August 202612September 2026 (to Sep 6)7
Tracked AI model releases by month, June to September 2026
ItemValue
June 20265
July 202621
August 202612
September 2026 (to Sep 6)7
Model releases by month from the release archive. September is counted through September 6 only and will grow; June through August are complete.Source: Capital & Compute AI model release archive

July was the outlier, not August. The GPT-5.6 family alone contributed three of those twenty-one, and Claude Opus 5, Kimi K3 and the Gemini 3.5 tiers filled most of the rest. August settled to twelve. September has already passed June’s whole-month total in three days, and it has done it with all four of the largest Western labs moving inside 72 hours of each other.

One caution on reading that last column: it counts three days, and two of the labs that release most often have not yet moved. Alibaba already has, with the Qwen3.8-Max-0902 snapshot; Z.ai and DeepSeek each shipped in the last week of August, and the Chinese labs in particular have run a cadence closer to fortnightly than monthly all year. Whether the month lands nearer twelve or nearer twenty depends mostly on them, and on whether Gemini 3.5 Pro arrives inside it. It has no date.

Did OpenAI release a new model in September 2026?

Yes. GPT-6 Astra was released on September 3, 2026 at $10 input and $50 output per million tokens, replacing GPT-5.6 Sol at the top of the lineup. It is the first increase in the frontier rate in years: Sol sits at $4 and $20, and those figures are themselves flagged promotional through at least November 21. Astra carries a 1,050,000-token context window, a five-step reasoning-effort dial and a knowledge cutoff of April 30, 2026, and it leads the independent Terminal-Bench 4.0 board at 58.18%.

Did Google release a new model in September 2026?

Yes. Gemini 3.8 Flash and the restricted Gemini 3.8 Flash Cyber arrived on September 2, at the same $0.75 and $3.75 per million tokens Gemini 3.7 Flash launched at on August 13. Google states on its own pricing page that this rate doubles to $1.50 and $7.50 on January 1, 2027, and the same expiry applies to 3.6 Flash and 3.7 Flash. Gemini 3.5 Pro, the top-end tier, is still announced with no date.

Which September model should you actually use?

It depends on when the work runs and on whether you are buying tokens or buying finished tasks.

For work shipping this year, Gemini 3.8 Flash is the cheapest capable option: $0.73 on a modeled multi-step agentic task against $1.12 for Muse Spark 1.3 and $7.68 for Fable 5.1, and the fastest at 306 output tokens per second. For work running past January 1, price it at the doubled rate of $1.50 and $7.50, at which point Muse Spark 1.3 is cheaper per task and scores two points higher on the independent index.

For hard agentic work, the ranking inverts against the rate card. GPT-6 Astra is the most expensive of the six per token and the cheapest per solved task on Terminal-Bench 4.0, at $11.88 at high effort against $32.69 for Fable 5.1 and $34.91 for Opus 5. Note that this is a measurement of a model inside a specific agent, Codex, and Fable 5.1 was measured inside Claude Code, so it is not a clean model-to-model comparison. Note too that Astra is not the most capable model by every measure: Artificial Analysis puts it at Intelligence Index 61 at max effort, behind Fable 5.1 at 66 and Opus 5 at 63.

Anthropic’s own guidance is to start with Claude Opus 5 and reach for Fable 5.1 only for demanding reasoning and long-horizon agentic work, or when Opus 5 at a higher effort setting still fails your evaluations. Opus 5 remains cheaper on every task profile that generates output.

The teams for whom September 1 was unambiguously good news are the ones already running Fable 5. For them this is roughly a fifth off the bill with no migration beyond three breaking API changes: forced tool use now returns an error, earlier models cannot read Fable 5.1 thinking blocks, and editing an earlier conversation turn invalidates them. The second one bites hardest on a mixed fleet, because a fallback path that drops out of Fable 5.1 into a cheaper model mid-conversation will break on the thinking history.

Current rates for every tracked model are in the AI model tracker, and the modeled cost comparison across all of them is in cost per coding task.

Frequently asked questions

What AI models were released in September 2026?
As of September 6, seven. Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 on September 1 at $10 and $50 per million tokens. Google released Gemini 3.8 Flash and the restricted Gemini 3.8 Flash Cyber on September 2, with 3.8 Flash at $0.75 and $3.75. Meta released Muse Spark 1.3 on September 2 at $1.25 and $4.25. Alibaba released the Qwen3.8-Max-0902 snapshot on September 2 at $2 and $6, unchanged from the August base. OpenAI released GPT-6 Astra on September 3 at $10 and $50. The first six held their predecessor headline rates; GPT-6 Astra raised the frontier rate 2.5 times over GPT-5.6 Sol.
Did Anthropic release a new model in September 2026?
Yes. Claude Fable 5.1 and Claude Mythos 5.1 were released on September 1, 2026. Fable 5.1 keeps the Fable 5 rates of $10 and $50 per million tokens and cuts the cache read from $1.00 to $0.25, which is worth about 21% on a long agentic task.
Did OpenAI or Google release a new model in September 2026?
Both did. Google released Gemini 3.8 Flash and the restricted Gemini 3.8 Flash Cyber on September 2. OpenAI released GPT-6 Astra on September 3 at $10 input and $50 output per million tokens, 2.5 times the GPT-5.6 Sol rate it replaces, and it leads the independent Terminal-Bench 4.0 leaderboard at 58.18 percent. Google has also announced Gemini 3.5 Pro as coming soon without a date; it is now the only unreleased model the tracker is watching.
How many AI models have been released in 2026 so far?
The release archive here tracks 45 dated launches from June 2026 onward: 5 in June, 21 in July, 12 in August and 7 in September through September 6. Every row carries the price the model launched at and a link to the provider primary source.
Why did Claude Sonnet 5 not get more expensive on September 1?
Anthropic cancelled the increase. Sonnet 5 launched at $2 and $10 per million tokens as introductory pricing billed as running to August 31, 2026, with a rise to $3 and $15 scheduled for September 1. Anthropic announced in August that the increase would not occur and the pricing documentation now lists $2 and $10 as the standard rate.

Sources

Anthropic (2026). Claude Fable 5.1 and Claude Mythos 5.1. Anthropic announcement. https://www.anthropic.com/claude-fable-and-mythos-5-1 Verified 2026-09-02.

Anthropic (2026). Pricing. Claude Platform documentation, including the Claude Sonnet 5 introductory-pricing note and the Fable 5.1 cache-multiplier footnote. https://platform.claude.com/docs/en/about-claude/pricing Verified 2026-09-02.

Anthropic (2026). Claude Fable 5.1. Claude Platform model documentation. https://platform.claude.com/docs/en/models/fable-5-1/overview Verified 2026-09-02.

Anthropic (2026). Models overview. Claude Platform documentation. https://platform.claude.com/docs/en/models/overview Verified 2026-09-02.

Google (2026). Introducing Gemini 3.8 Flash and 3.8 Flash Cyber. Google blog announcement. https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/ Verified 2026-09-03.

Google (2026). Gemini Developer API pricing. Google AI for Developers documentation. https://ai.google.dev/gemini-api/docs/pricing Verified 2026-09-03.

Meta AI Research (2026). Introducing Muse Spark 1.3. Meta AI Research blog. https://research.meta.ai/blog/introducing-muse-spark-1-3 Verified 2026-09-03.

Meta (2026). Model API pricing and rate limits. Meta Model API documentation. https://dev.meta.ai/docs/pricing-rate-limits/ Verified 2026-09-03.

Artificial Analysis (2026). Muse Spark 1.3: Meta reaches the frontier. Artificial Analysis, independent benchmarking service. https://artificialanalysis.ai/articles/muse-spark-1-3 Verified 2026-09-03.

Qwen (2026). Qwen3.8-Max-0902. QwenCloud model page, with per-token pricing and API reference. https://www.qwencloud.com/models/qwen3.8-max-0902 Verified 2026-09-06.

Qwen (2026). Model releases changelog. QwenCloud documentation, dating the 0902 snapshot to September 2, 2026. https://docs.qwencloud.com/changelog/models Verified 2026-09-06.

Alibaba Cloud (2026). Update notice for Qwen3.8-Max models. Model Studio notice: the qwen3.8-max endpoint transitions to qwen3.8-max-0902 effective September 5, 2026, billing unchanged. https://www.alibabacloud.com/en/notice/model_studio_update_notice_for_qwen38max_models_863 Verified 2026-09-06.

OpenAI (2026). API pricing. OpenAI developer documentation. https://developers.openai.com/api/docs/pricing Verified 2026-09-04.

OpenAI (2026). GPT-6 Astra. OpenAI developer documentation, model reference. https://developers.openai.com/api/docs/models/gpt-6-astra Verified 2026-09-04.

Terminal-Bench (2026). Terminal-Bench 4.0 leaderboard. Stanford and Laude Institute, independent benchmark. https://www.tbench.ai/leaderboard/terminal-bench/4.0 Verified 2026-09-04.

ARC Prize Foundation (2026). GPT-6 Astra on ARC-AGI-3. ARC Prize, independent benchmark. https://arcprize.org/blog/astra Verified 2026-09-04.

Monthly release counts are derived from the release archive maintained on this site, whose every row carries the provider primary source it was verified against.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to Models & benchmarks