AI Model Release Tracker
The latest and upcoming AI models in 2026, with release dates, per-token pricing, and a link to every provider's own page. Released models are verified against official pricing; upcoming ones appear only when a primary source confirms them. To turn a sticker rate into what a task actually costs, use thecost-per-task calculator or put two models head to head in the model comparison, and for monthly plan prices see the AI coding plan comparison.
What new AI models launched in 2026?
The most recent arrivals are OpenAI's GPT-6 Astra (September 3), Alibaba's Qwen3.8-Max-0902 (September 2), Anthropic's Claude Fable 5.1 (September 1), xAI's Grok 4.6 (August 12) and Alibaba's Qwen3.8-Max (August 3). The table below lists every tracked model with its current per-token rate, each read from the provider's own pricing page and dated.
For this month specifically, see the full roundup of new AI models released in September 2026. This page tracks what each model costs now; for the dated chronology of what shipped when, and the price each model launched at, see AI model releases by month.
Released and available
Models you can call today, newest first. Rates are USD per million tokens (Mtok), standard non-batch pricing, each verified against the provider's official page.
| Provider · Model | Status | Released / expected | Input / output (per Mtok) | Source |
|---|---|---|---|---|
OpenAI GPT-6 Astra | Sep 2026 | $10 / $50 | Source ↗ | |
Google Gemini 3.8 Flash | Sep 2026 | $0.75 / $3.75 | Source ↗ | |
Meta Muse Spark 1.3 | Sep 2026 | $1.25 / $4.25 | Source ↗ | |
Alibaba Qwen3.8-Max-0902 | Sep 2026 | $2 / $6 | Source ↗ | |
Anthropic Claude Fable 5.1 | Sep 2026 | $10 / $50 | Source ↗ | |
Zhipu GLM-5.3-Flash | Aug 2026 | $0.15 / $0.5 | Source ↗ | |
Alibaba Qwen3.8-27B | Aug 2026 | Not yet priced | Source ↗ | |
Google Gemini 3.7 Flash | Aug 2026 | $0.75 / $3.75 | Source ↗ | |
xAI Grok 4.6 | Aug 2026 | $2 / $6 | Source ↗ | |
Alibaba Qwen3.8-Max | Aug 2026 | $2 / $6 | Source ↗ | |
DeepSeek DeepSeek V4 Flash | Jul 2026 | $0.44 / $1.32 | Source ↗ | |
Anthropic Claude Opus 5 | Jul 2026 | $5 / $25 | Source ↗ | |
Moonshot Kimi K3 | Jul 2026 | $3 / $15 | Source ↗ | |
Thinking Machines Inkling | Jul 2026 | $1.87 / $4.68 | Source ↗ | |
Meta Muse Spark 1.1 | Jul 2026 | Not yet priced | Source ↗ | |
OpenAI GPT-5.6 Sol | Jul 2026 | $4 / $20 | Source ↗ | |
OpenAI GPT-5.6 Terra | Jul 2026 | $2 / $12 | Source ↗ | |
OpenAI GPT-5.6 Luna | Jul 2026 | $0.2 / $1.2 | Source ↗ | |
xAI Grok 4.5 | Jul 2026 | $2 / $6 | Source ↗ | |
Anthropic Claude Sonnet 5 | Jun 2026 | $2 / $10 | Source ↗ | |
Anthropic Claude Fable 5 | Jun 2026 | $10 / $50 | Source ↗ | |
Cohere North Mini Code | Jun 2026 | Not yet priced | Source ↗ | |
DeepSeek DeepSeek V4 Pro | Apr 2026 | $1.32 / $3.96 | Source ↗ | |
Anthropic Claude Opus 4.8 | – | $5 / $25 | Source ↗ | |
Anthropic Claude Sonnet 4.6 | – | $3 / $15 | Source ↗ | |
Anthropic Claude Haiku 4.5 | – | $1 / $5 | Source ↗ | |
OpenAI GPT-5.5 | – | $5 / $30 | Source ↗ | |
Google Gemini 3.1 Pro | – | $2 / $12 | Source ↗ | |
Google Gemini 3.5 Flash | – | $1.5 / $9 | Source ↗ | |
Google Gemini 3 Flash | – | $0.5 / $3 | Source ↗ | |
Alibaba Qwen3.7 Max | – | $2.5 / $7.5 | Source ↗ | |
Moonshot Kimi K2.7 Code | – | $0.95 / $4 | Source ↗ | |
Zhipu GLM-5.2 | – | $1.4 / $4.4 | Source ↗ | |
Thinking Machines Inkling-Small | – | Not yet priced | Source ↗ | |
Google Gemini 3.6 Flash | – | $0.75 / $3.75 | Source ↗ | |
Meta Muse Spark 1.2 | – | $1.25 / $4.25 | Source ↗ | |
Meta Muse Glimmer | – | Not yet priced | Source ↗ | |
Zhipu GLM-5.3 | – | Not yet priced | Source ↗ |
The rate card is only the starting point. Model any of these in the cost-per-task calculator to see what a real coding task costs, and why the cheapest per-token model is often not the cheapest to finish the job.
Upcoming AI models
The next wave is already taking shape. Every model below is either officially announced or surfaced by credible reporting, and each links to its source. The status badge says which is which, and no release date appears unless a provider has stated one. The day a model ships, its confirmed pricing and date move up to the tables above.
| Provider · Model | Status | Released / expected | Source |
|---|---|---|---|
Google Gemini 3.5 Pro | Coming soon | Source ↗ |
No price is recorded for a rumored model, however widely a figure is repeated. The most recent rumor to graduate was Claude Fable 5.1, carried here without a price through August and released on September 1, 2026: see what it actually shipped at against what the leak had claimed.
Retired or suspended
Models withdrawn, suspended, or superseded, including any pulled by a regulator. Kept for reference with the source that reported the change.
| Provider · Model | Status | Released / expected | Input / output (per Mtok) | Source |
|---|---|---|---|---|
Anthropic Claude Opus 4.1 | – | $15 / $75 | Source ↗ |
How this tracker stays honest
Every released model's per-token rate is read off the provider's own API pricing page, stored with that exact source URL, and dated to the day it was checked. A model reaches the upcoming table only when the provider has officially announced it or credible reporting has surfaced it, and that entry carries the reporting source. No release date is published unless a primary source states it, and no rumor appears without a citation. When a number cannot be confirmed against an official page, it is left off rather than guessed.
Released, preview, announced, rumored: what the labels mean
- Released: generally available, callable today, with a verified rate card.
- Preview: available to use but still rolling out or limited, so the rate may move.
- Announced: the provider has officially confirmed the model and, usually, a window, but it is not yet callable.
- Rumored: surfaced by credible reporting rather than the provider, with no confirmed date. Treated as a lead, attributed to its source, never as fact.
The sticker rate is not the bill
A low per-token price does not make a model cheap to run. What a coding task costs is set by how many tokens the agent burns reading context, reasoning, and generating output, and that swings by more than an order of magnitude between models and runs. A 2026 Microsoft Research preprint found the cheaper-listed model finished the same work at a higher cost in roughly a third of matchups, theprice reversal phenomenon. The honest unit is cost to finish a representative slice of your own tasks: thecost-per-task calculator models it for every row in this tracker, and the Claude Code cost breakdown walks the math end to end.
Frequently asked questions
What new AI models launched in 2026?
As of September 2026, the models tracked here include OpenAI's GPT-6 Astra (released September 3 at $10/$50 per Mtok, 2.5x the GPT-5.6 Sol rate it replaces, with a 2x surcharge on prompts over 272,000 tokens); Anthropic's Claude Fable 5.1 (released September 1 alongside the invitation-only Mythos 5.1, holding Fable 5's $10/$50 per Mtok and cutting only the cache read, from $1.00 to $0.25), Claude Opus 5 (released July 24), Sonnet 5 and Haiku 4.5, with Claude Opus 4.1 retired on the Claude API on August 5; OpenAI's GPT-5.5 and the GPT-5.6 family (Sol, Terra and Luna, generally available July 9, with Terra and Luna repriced July 30 and Sol taking its own promotional cut on August 21, held through at least November 21); xAI's Grok 4.6 (released August 12, 35 days after Grok 4.5); Meta's Muse Spark 1.1 and 1.2 plus Muse Glimmer (an Apache 2.0 open-weight model); Google's Gemini 3.7 Flash (released August 13 at an introductory rate that doubles on January 1, 2027), 3.6 Flash, 3.1 Pro, 3.5 Flash and 3 Flash; Thinking Machines' Inkling (released July 15, a 975B/41B-active MoE under Apache 2.0); and the leading China-built models DeepSeek V4 Pro and V4 Flash (both moving to peak and off-peak billing on August 16), Alibaba's Qwen3.7 Max, Qwen3.8-Max (released August 3, with open weights published August 13 under a custom licence rather than Apache 2.0) and the Qwen3.8-Max-0902 post-training snapshot (released September 2 at the same $2/$6 per Mtok), Moonshot's Kimi K3 (released July 16) and K2.7 Code, and Zhipu's GLM-5.2. Anthropic's Claude Fable 5 shipped June 9, 2026, was suspended June 12 under a US export-control directive, and returned in early July 2026 after the Commerce Department lifted the directive. Each row links to the provider's official page and is dated to when it was last verified.
What AI models are coming next?
As of September 2026 the tracker is watching one: Google's Gemini 3.5 Pro, which Google has announced and lists as "coming soon" with no date. The upcoming table emptied out this month. OpenAI's GPT-6, carried here as a rumor since June on an independent tracker alone, was released on September 3, 2026 as GPT-6 Astra and now sits in the released table at a verified $10 input and $50 output per million tokens. Anthropic's Claude Fable 5.1, carried as a rumor through August on community reports, was released on September 1. OpenAI's GPT-5.6 family, previously listed as expected, went generally available July 9. Each entry is attributed to its source. The tracker never publishes a release date a provider has not stated, and never lists a rumor without a citation.
How much do new AI models cost to use?
Per-token API rates span more than an order of magnitude, from GPT-5.6 Luna at $0.20 per million input tokens to Claude Fable 5, Fable 5.1 and GPT-6 Astra at $10. DeepSeek held the bottom of that range until 16:00 UTC on August 16, 2026, when it moved to peak and off-peak billing: V4 Flash input went from $0.14 flat to $0.22 off-peak and $0.44 at peak, so the rates shown here for both DeepSeek rows are the peak figures. But the sticker rate is not the bill: what a task actually costs is set by how many tokens an agent burns, and the cheapest per-token model is often not the cheapest to finish the work. GPT-6 Astra is the clearest case on this table. It costs 2.5 times GPT-5.6 Sol per token and matches Fable 5.1 exactly on the sticker, yet on the independent Terminal-Bench 4.0 board it finishes a task for $17.02 against the $32.69 Fable 5.1 costs, because it burns 1.53 billion tokens where Fable 5.1 burns 2.75 billion. Fable 5.1 makes the same point from the other direction: it matches Fable 5 on every headline rate and still costs about 21% less on a long agentic run, because the only number it changed was the cache read. Model any row in the cost-per-task calculator to see the real figure.
Where does this AI model data come from?
Every per-token rate is read off the provider's own official API pricing page, stored with that source URL, and stamped with the date it was checked. Release and announcement dates are recorded only when a primary source states them. Nothing here is taken from memory or an unsourced tracker.
Sources
Per-token rates for the released models are grounded to each provider's official API pricing page, verified September 2026:
Spotted a rate that has changed, or a model that is missing? Submit it or send a correction. Listings are free, editorial, and cannot be bought.
- Anthropic. (2026). Claude pricing (Fable 5, Opus 4.8, Sonnet 4.6, Haiku 4.5). claude.com/pricing
- OpenAI. (2026). API pricing (GPT-5.6). openai.com/api/pricing
- Google. (2026). Gemini Developer API pricing (Gemini 3.1 Pro, 3.5 Flash, 3 Flash). ai.google.dev/gemini-api/docs/pricing
- DeepSeek. (2026). API pricing (DeepSeek V4 Pro and V4 Flash). api-docs.deepseek.com/quick_start/pricing
- Alibaba Cloud. (2026). Model Studio model pricing (Qwen3.7 Max). alibabacloud.com/help/en/model-studio/models
- QwenCloud. (2026). Qwen3.8-Max-0902 model page and pricing (same $2/$6 rate as Qwen3.8-Max). qwencloud.com/models/qwen3.8-max-0902
- Moonshot AI. (2026). Kimi API pricing (Kimi K3 and K2.7 Code). platform.kimi.ai/docs/pricing
- Zhipu AI (via OpenRouter). (2026). GLM-5.2 API pricing (Z.ai's own rate card still rolling out). openrouter.ai/z-ai/glm-5.2
Latest analysis
- Best Local Models for Hermes Agent in 2026
Hermes Agent rejects any model under 64,000 tokens of context. Which open-weight models clear that bar on 8GB, 16GB, 24GB and 32GB of VRAM.
- GPT-6 Astra: Pricing, Benchmarks, Cost
GPT-6 Astra lists $10 and $50 per million tokens, 2.5x GPT-5.6 Sol. On Terminal-Bench it still solves a task for half of what Claude Fable 5.1 costs.
- Gemini 3.8 Flash and Muse Spark 1.3 Pricing
Gemini 3.8 Flash costs $0.75 per million input tokens until December 31, then doubles. Muse Spark 1.3 is $1.25, or $0.10 if Meta can train on you.
- Claude Fable 5.1: Pricing, Benchmarks, Cost
Claude Fable 5.1 holds Fable 5 rates at $10/$50 per million tokens and cuts the cache read to $0.25. Benchmarks, specs, and modeled cost per task.