AI model releases, month by month
A dated record of every AI model release we can verify against the lab that shipped it, grouped by the month it happened. 56 releases so far, each with the price it launched at, the context window, whether the weights are open, and a link to the announcement. For what these models cost today and what is expected next, see the model tracker; for how they rank on independent benchmarks, the value leaderboard.
How many AI models were released in September 2026?
18 distinct model releases in September 2026 are verifiable against a dated source, from 8 labs. 8 shipped with open weights. The flagship launches were DeepSeek V4.1 Flash, GPT-6 Astra, K2 Horizon 375B-A23B, Gemini 3.8 Flash, Muse Spark 1.3, Claude Fable 5.1. That number is smaller than the counts on auto-ingested catalogs, which list every fine-tune and re-hosted variant; counted here is a distinct model a lab announced as a release.
What is the most recent AI model release?
DeepSeek V4.1 Flash from DeepSeek, released Sep 10, 2026. September's only price cut: peak input falls 32% and the cache read 57% against V4 Flash. A 552B causal encoder-decoder activating 8B on prefill, MIT weights. It launched at $0.3 / $1.2 per million tokens.
| Day of September 2026 | Releases | Models |
|---|---|---|
| 1 | 2 | Claude Fable 5.1, Claude Mythos 5.1 |
| 2 | 4 | Gemini 3.8 Flash, Muse Spark 1.3, Gemini 3.8 Flash Cyber, Qwen3.8-Max-0902 |
| 3 | 7 | GPT-6 Astra, K2 Horizon 375B-A23B, K2 Horizon 0.9B, K2 Horizon 36B-A4B, K2 Horizon 3.7B, K2 Horizon 32B, K2 Horizon 7B |
| 4 | 1 | Ling-3.0-flash-Sante |
| 8 | 2 | GPT Image 2.5 Flare, GPT Image 2.5 Sunburst |
| 9 | 1 | Ling-3.0-flash-Fin |
| 10 | 1 | DeepSeek V4.1 Flash |
Releases do not arrive evenly. September 2026 produced 18 of them across just 7 separate days, and the most any one day carried was 7, on September 3. The pattern is competitive rather than coincidental: labs hold a launch until a rival moves, then answer within hours, which is why the calendar above is a run of spikes separated by quiet weeks rather than a steady drip.
| Item | Value |
|---|---|
| DeepSeek V4.1 Flash | $1.2 |
| Gemini 3.8 Flash | $3.75 |
| Muse Spark 1.3 | $4.25 |
| Qwen3.8-Max-0902 | $6 |
| GPT Image 2.5 Flare | $30 |
| GPT Image 2.5 Sunburst | $30 |
| GPT-6 Astra | $50 |
| Claude Fable 5.1 | $50 |
| Claude Mythos 5.1 | $50 |
The spread inside one month is about 42x, from DeepSeek V4.1 Flash at $1.2 per million output tokens to Claude Mythos 5.1 at $50. Read that as a range of positions, not a quality ranking: the cheapest per-token model is frequently not the cheapest way to finish a task, because a model that reasons for four times as many tokens erases a four-times-lower rate. The cost-per-task ranking works the figure the other way round, from a completed task backwards.
Has any AI model raised its price after launch?
No. Not one of the 23 releases we can price both ways has raised a published per-token rate since the day it shipped. 18 still charge exactly what they charged on release day. 4 have cut, some of them steeply. The direction is one-way, and that is the single most useful thing a release record can tell you that a list of dates cannot.
This comparison exists because two datasets are kept apart on purpose. This page freezes the rate a lab published on launch day and never updates it; the model tracker overwrites with the current rate. Holding them together is what turns a timeline into a price history. A current-price table alone cannot show it, because a model repriced downward six months later simply looks cheap.
| Item | At launch | Today |
|---|---|---|
| GPT-5.6 Luna | $6 | $1.2 |
| Gemini 3.6 Flash | $7.5 | $3.75 |
| GPT-5.6 Sol | $30 | $20 |
| GPT-5.6 Terra | $15 | $12 |
The cuts cluster in one place. 3 of the 4 are OpenAI repricing the GPT-5.6 family after GPT-6 Astra arrived above it, and the deepest, GPT-5.6 Luna at -80%, is the cheapest tier of that family. Anthropic has not moved a single rate in this record: Claude Fable 5, Fable 5.1, Opus 5 and Sonnet 5 all still list what they listed on launch day. Read that as a pricing strategy rather than a coincidence. The labs that cut are the ones defending a workhorse tier against cheaper rivals; the lab that holds is selling on capability.
One row needs its own caveat, and it is the only apparent increase on the page. DeepSeek V4 Flash 0731 looks like a rise from $0.28 to $1.32 per million output tokens, but the billing model changed rather than the rate: DeepSeek moved to peak and off-peak pricing on 2026-08-16, so the launch figure and the current figure are not the same kind of number. It is excluded from the cut count above rather than printed as a percentage, because the percentage would be arithmetic on two different things.
| Model | At launch | Today | Change, output |
|---|---|---|---|
| DeepSeek V4.1 FlashDeepSeek | $0.3 / $1.2 | $0.3 / $1.2 | No change |
| GPT-6 AstraOpenAI | $10 / $50 | $10 / $50 | No change |
| Gemini 3.8 FlashGoogle | $0.75 / $3.75 | $0.75 / $3.75 | No change |
| Muse Spark 1.3Meta | $1.25 / $4.25 | $1.25 / $4.25 | No change |
| Qwen3.8-Max-0902Alibaba | $2 / $6 | $2 / $6 | No change |
| Claude Fable 5.1Anthropic | $10 / $50 | $10 / $50 | No change |
| GLM-5.3-FlashZhipu | $0.15 / $0.5 | $0.15 / $0.5 | No change |
| Gemini 3.7 FlashGoogle | $0.75 / $3.75 | $0.75 / $3.75 | No change |
| Grok 4.6xAI | $2 / $6 | $2 / $6 | No change |
| Muse Spark 1.2Meta | $1.25 / $4.25 | $1.25 / $4.25 | No change |
| Qwen3.8-MaxAlibaba | $2 / $6 | $2 / $6 | No change |
| DeepSeek V4 Flash 0731DeepSeek | $0.14 / $0.28 | $0.44 / $1.32 | Billing change |
| Claude Opus 5Anthropic | $5 / $25 | $5 / $25 | No change |
| Gemini 3.6 FlashGoogle | $1.5 / $7.5 | $0.75 / $3.75 | -50% |
| Kimi K3Moonshot | $3 / $15 | $3 / $15 | No change |
| InklingThinking Machines | $1.87 / $4.68 | $1.87 / $4.68 | No change |
| GPT-5.6 SolOpenAI | $5 / $30 | $4 / $20 | -33% |
| GPT-5.6 LunaOpenAI | $1 / $6 | $0.2 / $1.2 | -80% |
| GPT-5.6 TerraOpenAI | $2.5 / $15 | $2 / $12 | -20% |
| Grok 4.5xAI | $2 / $6 | $2 / $6 | No change |
| Claude Sonnet 5Anthropic | $2 / $10 | $2 / $10 | No change |
| GLM-5.2Zhipu | $1.4 / $4.4 | $1.4 / $4.4 | No change |
| Claude Fable 5Anthropic | $10 / $50 | $10 / $50 | No change |
Why do AI release trackers disagree on dates?
Because most of them ingest an API catalog instead of reading an announcement, so they record when a model appeared in a listing rather than when the lab shipped it. Two live examples from September 2026 alone. Grok STT 1.0 is dated July 23, 2026 on one major timeline; xAI actually released it on March 16, 2026, and July 23 is the day it was added to OpenRouter. Kimi K3 is dated July 14 by one tracker and July 16 by two others; Moonshot unveiled it on July 16. Every date below is read off the lab announcement, and the 9 rows that rest on reporting are marked as such.
How this tracker stays honest
AI models released in September 2026
18 releases from 8 labs, 8 of them with open weights. Launch prices ran from $1.2 to $50 per million output tokens across the 9 of 18 rows that shipped with a published rate card.
September moved prices in both directions, which is why a one-line summary of it is always wrong. Anthropic, Google, Meta and Alibaba all shipped without touching a rate card. Then GPT-6 Astra listed at $10 and $50 per million tokens, two and a half times what GPT-5.6 Sol charges and the first increase at the frontier tier in years, and six days later DeepSeek V4.1 Flash cut the efficiency tier, taking peak input 32% and the cache read 57% below the V4 Flash it retires. Read the rise against token consumption rather than on its own: Astra burns far fewer tokens per task than its price implies. Two releases shipped with MIT weights, and one of them is the cut. September also delivered the largest fully open release in the record: the Abu Dhabi lab IFM shipped six Apache 2.0 models, 0.9B to 375B, with training data and checkpoints, and none carries a first-party per-token rate.
The flagship launches were DeepSeek V4.1 Flash, GPT-6 Astra, K2 Horizon 375B-A23B, Gemini 3.8 Flash, Muse Spark 1.3, Claude Fable 5.1. Read the September analysis, or filter the full record to this month in the table below.
AI models released in August 2026
12 releases from 7 labs, 5 of them with open weights. Launch prices ran from $0.177 to $6 per million output tokens across the 10 of 12 rows that shipped with a published rate card.
August was the open-weight month. Five of the twelve releases shipped with downloadable weights, and four of them came from Chinese labs, with Alibaba alone accounting for three. The frontier barely moved: Grok 4.6 took the top intelligence score at 61, matching what Claude Opus 5 scored in July rather than beating it. Prices clustered low, from $0.177 to $6 per million output tokens, because the volume was in workhorse and open-weight tiers rather than at the top.
The flagship launches were GLM-5.3-Flash, Qwen3.8-Flash-Next, GLM-5.3, Qwen3.8-27B, Gemini 3.7 Flash, Grok 4.6, Muse Glimmer, Qwen3.8-Max. Read the August analysis, or filter the full record to this month in the table below.
AI models released in July 2026
21 releases from 14 labs, 4 of them with open weights. Launch prices ran from $0.2 to $30 per million output tokens across the 11 of 21 rows that shipped with a published rate card.
July was the busiest month in the record by a wide margin: 21 releases from 14 different labs, four times the June count. It is also the month with the widest price spread, $0.20 to $30 per million output tokens, and the month where the launch record starts diverging from the current one. Three of the five per-token price cuts on this page are July models. Ten of the 21 rows carry no per-token rate at all, because speech, image and video models are priced by the second, the character or the image.
The flagship launches were DeepSeek V4 Flash 0731, Claude Opus 5, Kimi K3, GPT-5.6 Sol, Grok 4.5. Read the July analysis, or filter the full record to this month in the table below.
AI models released in June 2026
5 releases from 3 labs, 2 of them with open weights. Launch prices ran from $4.4 to $50 per million output tokens across the 4 of 5 rows that shipped with a published rate card.
June is a partial month: this record starts on June 9, 2026, so earlier June releases are not covered and the five rows here are not a complete count. What they do show is a quiet baseline before the July surge, with three labs shipping and Anthropic accounting for three of the five.
The flagship launches were Claude Sonnet 5, GLM-5.2, Claude Fable 5.
Filter every tracked release
Search by model, lab or description, then narrow by category or lab. Price is the rate the model launched at, not necessarily what it costs today.
| Released | Model | Launch price | AA Index | Context | Source |
|---|---|---|---|---|---|
| Sep 10, 2026 | DeepSeek V4.1 Flash DeepSeek· open weight(MIT) September's only price cut: peak input falls 32% and the cache read 57% against V4 Flash. A 552B causal encoder-decoder activating 8B on prefill, MIT weights. | $0.3 / $1.2 | – | 1M | Official |
| Sep 9, 2026 | Ling-3.0-flash-Fin Ant Group· open weight(MIT) Finance tuning of Ling-3.0-flash, 124B total and 5.1B active, MIT weights. Ant Group publishes no per-token rate of its own; hosting is third-party. | – | – | 262K | Official |
| Sep 8, 2026 | GPT Image 2.5 Flare OpenAI The faster default of the two GPT Image 2.5 API variants. Billed per token, not per image, which is why it carries a rate where earlier image rows here carry none. | $8 / $30 | – | – | Official |
| Sep 8, 2026 | GPT Image 2.5 Sunburst OpenAI The precision-control GPT Image 2.5 variant, slower to generate, on the same $8 and $30 per million token rate as Flare. Text input is billed separately at $5. | $8 / $30 | – | – | Official |
| Sep 4, 2026 | Ling-3.0-flash-Sante Ant Group Health and medicine tuning of Ling-3.0-flash, 124B total and 5.1B active. Hosted API only: no weights and no own rate card, unlike the Fin sibling. | – | – | 262K | Reporting |
| Sep 3, 2026 | GPT-6 Astra OpenAI OpenAI's new flagship at $10/$50 per Mtok, 2.5x GPT-5.6 Sol and level with Claude Fable 5.1. Tops Terminal-Bench 4.0 at 58.18% while burning 1.53B tokens to Fable 5.1's 2.75B. | $10 / $50 | 61 | 1.05M | Official |
| Sep 3, 2026 | K2 Horizon 375B-A23B IFM· open weight(Apache 2.0) 375B MoE activating 23B per token, Apache 2.0 with training data and checkpoints. AA 47 at launch; Terminal-Bench 2.1 70.2% self-audited down to 66.9%. | – | 47 | 524K | Official |
| Sep 3, 2026 | K2 Horizon 0.9B IFM· open weight(Apache 2.0) 0.9B watch-scale dense model, 128K context. AIME 2026 48.5 and HumanEval+ 79.9 per IFM, Apache 2.0 weights. | – | – | 131K | Official |
| Sep 3, 2026 | K2 Horizon 36B-A4B IFM· open weight(Apache 2.0) 36B total with 4B active via MoVA sparsity inside attention. Near-dense-32B results at one-eighth the active parameters, 524K context. | – | – | 524K | Official |
| Sep 3, 2026 | K2 Horizon 3.7B IFM· open weight(Apache 2.0) 3.7B model trained on the same 22T tokens as the 7B and 32B. SWE-bench Verified 68.6 vs 41.2 for Qwen3.5-4B on IFM's table. | – | – | 524K | Official |
| Sep 3, 2026 | K2 Horizon 32B IFM· open weight(Apache 2.0) Dense 32B for local serving, 524K context. IFM's own tables trail Qwen3.8-27B on coding; ships Apache 2.0 with full training records. | – | – | 524K | Official |
| Sep 3, 2026 | K2 Horizon 7B IFM· open weight(Apache 2.0) 7B phone-scale model: SWE-bench Verified 70.6, Terminal-Bench 39.1 per IFM. IFM disclosed and excluded a contaminated 82-point SWE-bench run. | – | – | 524K | Official |
| Sep 2, 2026 | Gemini 3.8 Flash Google Third straight Flash generation at $0.75/$3.75 per Mtok, and the third sharing one expiry: the rate doubles on 2027-01-01. AA Intelligence Index 59 at high effort, against 56 for 3.7 Flash. | $0.75 / $3.75 | 59 | 1M | Official |
| Sep 2, 2026 | Muse Spark 1.3 Meta Meta held the 1.2 rate card exactly at $1.25/$4.25 per Mtok and moved AA Intelligence Index 57 to 61 at xhigh. The $0.10/$0.20 contributor tier buys its discount with training rights. | $1.25 / $4.25 | 61 | 1.05M | Official |
| Sep 2, 2026 | Gemini 3.8 Flash Cyber Google Vulnerability-hunting twin of 3.8 Flash, restricted to the Fairwind Program for governments, critical infrastructure operators and software maintainers. No published token rate. | – | – | – | Official |
| Sep 2, 2026 | Qwen3.8-Max-0902 Alibaba Post-trained refresh of Qwen3.8-Max at the same $2/$6: Terminal-Bench 3.0 doubles to 29.0 on Qwen's own table. API-only, no new weights. | $2 / $6 | – | 1M | Official |
| Sep 1, 2026 | Claude Fable 5.1 Anthropic Held Fable 5's $10/$50 per Mtok and cut only the cache read, $1.00 to $0.25, a 0.025x multiplier that undercuts Opus 5's cache rate at twice the base input price. | $10 / $50 | – | 1M | Official |
| Sep 1, 2026 | Claude Mythos 5.1 Anthropic The same model as Fable 5.1 with safeguards lifted in some areas, at the same $10/$50 and the same $0.25 cache read. Project Glasswing invitation only, never generally available. | $10 / $50 | – | 1M | Official |
| Aug 26, 2026 | GLM-5.3-Flash Zhipu· open weight(MIT) Z.ai's reveal of the Ox Alpha stealth model: a 320B-total, 18B-active MoE, MIT-licensed with weights on Hugging Face, scoring 57 on the Artificial Analysis Intelligence Index. | $0.15 / $0.5 | 57 | 1.05M | Official |
| Aug 26, 2026 | Qwen3.8-Flash-Next Alibaba· open weight(Qwen Community 1.0) An open-weight preview of the Qwen4 architecture: 125B total, 6B active, plus a 51B n-gram embedding layer. Multimodal in, text out, 262K native context extensible to 1M. | $0.15 / $0.47 | – | 262K | Official |
| Aug 21, 2026 | DeepSeek V4 Flash Vision Exp DeepSeek DeepSeek's first multimodal vision model, experimental, priced exactly as V4 Flash: $0.44/$1.32 per Mtok at peak and half that off-peak. Images bill as input tokens by dimension. | $0.44 / $1.32 | – | 1.05M | Official |
| Aug 20, 2026 | Hy-MT2-30B-A3B Tencent Translation flagship, about 3B active params, 33 language pairs plus five Chinese dialect pairs. An 8K context window is the point: a tool, not an agent. | $0.074 / $0.295 | – | 8K | Reporting |
| Aug 20, 2026 | Hy-MT2-1.8B Tencent Compact sibling of Hy-MT2-30B-A3B with the same language coverage, at the cheapest published output rate of any model released in August 2026. | $0.044 / $0.177 | – | 8K | Reporting |
| Aug 14, 2026 | GLM-5.3 Zhipu Same base as GLM-5.2, all gains from post-training; Terminal-Bench 3.0 jumps 4.6 to 28.3. Unpriced at launch, then listed at the GLM-5.2 rate. Weights still held back. | $1.4 / $4.4 | – | 1.05M | Official |
| Aug 14, 2026 | Qwen3.8-27B Alibaba· open weight(Apache 2.0) The smaller Qwen3.8 sibling, delivered on the promised week: dense 27B native vision-language, Apache 2.0 with no revenue gate, unlike Qwen3.8-Max. | – | – | 262K | Official |
| Aug 13, 2026 | Gemini 3.7 Flash Google Launched at an introductory $0.75/$3.75 that Google states doubles on 1 January 2027, with DeepSWE v1.1 up to 65.3% from 3.6 Flash 49.0%. | $0.75 / $3.75 | 56 | – | Official |
| Aug 12, 2026 | Grok 4.6 xAI Arrived 35 days after Grok 4.5 at the same $2/$6, but with cached input raised from $0.30 to $0.50 per Mtok. Ties GPT-5.6 Sol at 61 on the AA index. | $2 / $6 | 61 | 500K | Official |
| Aug 10, 2026 | Muse Glimmer Meta· open weight(Apache 2.0) Meta returns to a plain Apache 2.0 licence with a 29.6B dense multimodal model that fits one consumer GPU at 4-bit. Built for local agents, not the leaderboard. | – | 35 | 131K | Reporting |
| Aug 5, 2026 | Muse Spark 1.2 Meta Meta holds the 1.1 rate of $1.25/$4.25 and adds a $0.10/$0.20 contributor tier, plus Muse Code, a terminal coding agent, in beta. | $1.25 / $4.25 | 57 | 1M | Reporting |
| Aug 3, 2026 | Qwen3.8-Max Alibaba· open weight(qwen3.8-max (custom, revenue-gated)) The first Max-class Qwen with published weights (13 August) and, at AA 58, the highest-scoring open-weights model. The licence is custom, not Apache 2.0. | $2 / $6 | 58 | 262K | Official |
| Jul 31, 2026 | DeepSeek V4 Flash 0731 DeepSeek A re-post-trained V4 Flash at the same $0.14/$0.28 price, scoring 50 on the Artificial Analysis Intelligence Index, 6 points above the pricier V4 Pro. API public beta; 0731 weights not yet posted. | $0.14 / $0.28 | 50 | 1.05M | Official |
| Jul 24, 2026 | Claude Opus 5 Anthropic Took the top spot on the Artificial Analysis Intelligence Index at 61, and did it at half the per-token price of Claude Fable 5. | $5 / $25 | 61 | 1M | Official |
| Jul 23, 2026 | FLUX 3 Black Forest Labs One network spanning image, video, audio and robot action prediction. Video runs in early access; the open-weight Dev build is promised later in 2026. | – | – | – | Official |
| Jul 23, 2026 | Ling-3.0-flash Ant Group A 124B mixture-of-experts firing only 5.1B parameters per token, which Ant says matches its own 1T flagship. Announced as open weight, but shipped API-only with the weights still unpublished. | – | – | 256K | Official |
| Jul 21, 2026 | Gemini 3.6 Flash Google Google's workhorse tier, cut to $7.50 output from the $9.00 that Gemini 3.5 Flash charged, on a claimed 17% drop in output tokens per task. | $1.5 / $7.5 | – | 1.05M | Official |
| Jul 21, 2026 | Laguna S 2.1 poolside· open weight(OpenMDW-1.1) 118B total and 8B active, released under OpenMDW-1.1 at $0.10 in and $0.20 out, the cheapest agentic coding model to ship in July. | $0.1 / $0.2 | – | 1M | Official |
| Jul 21, 2026 | Gemini 3.5 Flash Cyber Google Tuned to find and patch software vulnerabilities, and restricted to governments and trusted partners through the CodeMender pilot. | – | – | – | Official |
| Jul 21, 2026 | Gemini 3.5 Flash-Lite Google The cheapest launch price of any July model at $0.30 in, aimed at classification, extraction and routing rather than reasoning. | $0.3 / $2.5 | – | 1.05M | Official |
| Jul 21, 2026 | Qwen-Image-3.0 Alibaba Closed and invite-only at launch, with no model card, weights or published benchmarks to check the multi-panel layout claims against. | – | – | – | Reporting |
| Jul 20, 2026 | Qwen-Audio-3.0-TTS Plus Alibaba Took first place on the Artificial Analysis text-to-speech arena. Billed per character, not per token, at roughly a third of what ElevenLabs charges. | – | – | – | Reporting |
| Jul 20, 2026 | Qwen-Audio-3.0-TTS Flash Alibaba The real-time tier of the same text-to-speech release, covering 16 languages and 20 Chinese dialect regions, hosted only on Alibaba Cloud Model Studio. | – | – | – | Reporting |
| Jul 19, 2026 | Qwen3.8-Max-Preview Alibaba A 2.4T-parameter preview shown at the World AI Conference with no model card, no license and no independent benchmarks published alongside it. | – | – | 1M | Reporting |
| Jul 16, 2026 | Kimi K3 Moonshot· open weight 2.8T parameters and the highest-scoring open-weight model of the month. Weights followed on July 26, a day inside Moonshot's own deadline. | $3 / $15 | 57 | 1.05M | Official |
| Jul 15, 2026 | Inkling Thinking Machines· open weight(Apache 2.0) 975B total and 41B active under Apache 2.0, the most permissive license anyone has attached to a model at that parameter scale. | $1.87 / $4.68 | 41 | 1M | Official |
| Jul 9, 2026 | GPT-5.6 Sol OpenAI The reasoning tier of the GPT-5.6 line. Launched at $5/$30, then cut to $4/$20 on Aug 21, 2026 (promotional through at least Nov 21) after being excluded from the July 30 Terra and Luna cuts. | $5 / $30 | 59 | 1.05M | Official |
| Jul 9, 2026 | GPT-5.6 Luna OpenAI The high-volume tier, launched at $1 in and $6 out, the best intelligence-per-dollar OpenAI shipped in July. Cut 80% on July 30 to $0.20/$1.20, under the retired GPT-5.4 nano floor. | $1 / $6 | 51 | 1.05M | Official |
| Jul 9, 2026 | GPT-5.6 Terra OpenAI The middle tier, launched at half of Sol and within four AA index points of it. Cut 20% on July 30, 2026 to $2/$12, half of Sol's promotional $4/$20. | $2.5 / $15 | 55 | 1.05M | Official |
| Jul 9, 2026 | Muse Spark 1.1 Meta Meta Superintelligence Labs shipped its first paid model, breaking the open-weight-only posture Meta had held since Llama. | – | – | 1M | Official |
| Jul 8, 2026 | Grok 4.5 xAI Trained on real Cursor session data and priced at $2 in and $6 out, well under half of what Opus 4.8 and GPT-5.5 charged at the time. | $2 / $6 | 54 | 500K | Official |
| Jul 8, 2026 | SWE-1.7 Cognition Post-trained on top of Moonshot's already RL-heavy Kimi K2.7 Code, and served inside Devin only. Not sold as a standalone API. | – | – | – | Official |
| Jul 7, 2026 | Cohere Transcribe Arabic Cohere· open weight(Apache 2.0) A 2B open-weight speech recognition model built for Arabic dialect variation and Arabic-English code-switching. | – | – | – | Official |
| Jun 30, 2026 | Claude Sonnet 5 Anthropic Anthropic's mid-tier Claude at a $2/$10 introductory rate billed as expiring on August 31. Anthropic cancelled that rise in August and made $2/$10 the standard price. | $2 / $10 | – | 1M | Official |
| Jun 16, 2026 | GLM-5.2 Zhipu· open weight(MIT) Z.ai's coding flagship, MIT-licensed with weights on Hugging Face, at $1.40/$4.40 per Mtok. The same rate GLM-5.1 carried and the same rate GLM-5.3 would carry three months later. | $1.4 / $4.4 | – | – | Official |
| Jun 9, 2026 | Claude Fable 5 Anthropic Anthropic's Mythos-class flagship made safe for general use, at $10/$50 per Mtok. Suspended three days later by a US export-control directive, restored after that order lifted on June 30. | $10 / $50 | – | 1M | Official |
| Jun 9, 2026 | Claude Mythos 5 Anthropic The same underlying model as Fable 5 with safeguards lifted in some areas, at the same $10/$50. Restricted to Project Glasswing partners and select researchers, never generally available. | $10 / $50 | – | – | Official |
| Jun 9, 2026 | North Mini Code Cohere· open weight(Apache 2.0) Cohere's first developer model: a 30B-total, 3B-active sparse MoE coding agent, Apache 2.0, running on one H100 at FP8. Free on hosted endpoints, so it launched with no rate card. | – | – | 262K | Official |
Launch price is the standard non-batch rate the lab published on release day, in USD per million tokens. A dash means the model has no per-token rate card: speech and image models are billed per character or per generation, and some models ship only inside a product. AA Index is the Artificial Analysis Intelligence Index at or near launch. Source links to the lab announcement, or is marked Reporting where no primary source exists.
What the September record actually shows
Three things stand out once the month is laid out chronologically rather than as a leaderboard.
Prices moved in both directions in the same ten days. GPT-6 Astra listed at $10 and $50 per million tokens, two and a half times what GPT-5.6 Sol charges and the first increase at the frontier tier in years. Six days later DeepSeek V4.1 Flash cut the efficiency tier, taking peak input 32% and the cache read 57% below the V4 Flash it retires. Every other priced release this month held its predecessor's headline rate exactly. A month summarised as either a rise or a cut is summarised wrongly. Put any two of these against the same task with the model comparison tool.
The open-weight releases were the ones that changed a price. 8 of 18 releases shipped with downloadable weights, six of them Apache 2.0 from IFM's K2 fleet and two MIT. That is no longer a small share, but the cut above still came from one of the MIT pair rather than from a proprietary lab defending share, and the other, Ant Group's finance-tuned Ling-3.0-flash-Fin, ships weights while publishing no rate card of its own at all, a position all six K2 models share. The open-weight tier is now setting the price floor rather than following it.
Most of the month was not frontier models. Five of the 18 releases were frontier launches; the rest were workhorse, specialist, image and open-weight models, including two domain fine-tunes aimed at finance and at medicine, a vulnerability-hunting Gemini variant nobody outside a vetted programme can call, and a six-model open fleet with no rate card at all. That is the shape of a maturing market: the frontier gets a few launches a month and the volume moves to specialised models that do one job cheaply. Nine rows this month carry no per-token rate at all, which is why they show a dash rather than a number.
For the narrative version of this month, with the benchmark detail and what it means for what you pay, read the September 2026 model roundup. For which of these you can call at no cost, see free AI models, and for what the benchmark names in each announcement actually measure, the AI benchmark directory.
Frequently asked questions
- How many AI models were released in September 2026?
- 18 model releases in September 2026 are verifiable against a dated source, from 8 different labs. That count is deliberately narrower than the auto-ingested catalogs, which list every fine-tune, quantization and re-hosted variant and run to dozens of rows a month. Counted here are distinct models a lab announced as a release: DeepSeek V4.1 Flash, GPT-6 Astra, K2 Horizon 375B-A23B, Gemini 3.8 Flash, Muse Spark 1.3, Claude Fable 5.1 were the flagship launches, and 8 of the 18 shipped with open weights.
- What was the most recent AI model release?
- DeepSeek V4.1 Flash from DeepSeek, released Sep 10, 2026. September's only price cut: peak input falls 32% and the cache read 57% against V4 Flash. A 552B causal encoder-decoder activating 8B on prefill, MIT weights.
- Why do AI model release trackers disagree on dates?
- Because most of them ingest a catalog rather than read an announcement, so they record the date a model appeared in an API listing rather than the date the lab shipped it. Two live examples: Grok STT 1.0 is dated July 23, 2026 by one major timeline, but xAI released it on March 16, 2026 and July 23 is only when it was added to OpenRouter. Kimi K3 is dated July 14 by one tracker and July 16 by two others; Moonshot unveiled it on July 16. Every date on this page is read off the lab announcement, and the handful that are not are labelled as reporting.
- What is a launch price, and why track it separately?
- A launch price is the standard per-token rate the lab published on release day. It is worth freezing because per-token rates move: a model repriced downward six months later looks cheap in a current price table, which hides how the market actually shifted. Holding the launch price next to the current rate is how you see the drift. No other release tracker records it.
- Which lab shipped the most models in September 2026?
- IFM, with 6. The full split: IFM 6, OpenAI 3, Ant Group 2, Anthropic 2, Google 2, Alibaba 1, DeepSeek 1, Meta 1. Volume is not quality: a three-model drop in one announcement is one product decision, not three independent launches.
- How much did AI model prices vary at launch?
- By about 42x within a single month. In September 2026 the cheapest launch output rate was $1.2 per million tokens (DeepSeek V4.1 Flash) and the most expensive was $50 (Claude Mythos 5.1). Per-token price is a poor guide to what work costs, though: token consumption varies more between models than the sticker rate does.
- Where does this release data come from?
- Each row is read off the lab's own announcement, documentation or press release and stamped with the date it was checked. 47 of 56 rows are sourced that way. 9 rest on reporting because the lab published no announcement page, and those are marked as such in the table and in the JSON. Nothing is taken from memory or from another tracker.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.
Sources
One primary source per release, with the date it was checked. 47 of 56cite the lab's own announcement, documentation or press release; rows marked reporting had no citable announcement from the lab. The full list is collapsed because it runs to 56 entries.
Show all 56 sources
- DeepSeek (2026). DeepSeek V4.1 Flash. Announcement. Released Sep 10, verified Sep 11.
- Ant Group (2026). Ling-3.0-flash-Fin. Announcement. Released Sep 9, verified Sep 11.
- OpenAI (2026). GPT Image 2.5 Flare. Announcement. Released Sep 8, verified Sep 11.
- OpenAI (2026). GPT Image 2.5 Sunburst. Announcement. Released Sep 8, verified Sep 11.
- Ant Group (2026). Ling-3.0-flash-Sante. Reporting. Released Sep 4, verified Sep 11.
- OpenAI (2026). GPT-6 Astra. Announcement. Released Sep 3, verified Sep 4.
- IFM (2026). K2 Horizon 375B-A23B. Announcement. Released Sep 3, verified Sep 13.
- IFM (2026). K2 Horizon 0.9B. Announcement. Released Sep 3, verified Sep 13.
- IFM (2026). K2 Horizon 36B-A4B. Announcement. Released Sep 3, verified Sep 13.
- IFM (2026). K2 Horizon 3.7B. Announcement. Released Sep 3, verified Sep 13.
- IFM (2026). K2 Horizon 32B. Announcement. Released Sep 3, verified Sep 13.
- IFM (2026). K2 Horizon 7B. Announcement. Released Sep 3, verified Sep 13.
- Google (2026). Gemini 3.8 Flash. Announcement. Released Sep 2, verified Sep 3.
- Meta (2026). Muse Spark 1.3. Announcement. Released Sep 2, verified Sep 3.
- Google (2026). Gemini 3.8 Flash Cyber. Announcement. Released Sep 2, verified Sep 3.
- Alibaba (2026). Qwen3.8-Max-0902. Announcement. Released Sep 2, verified Sep 6.
- Anthropic (2026). Claude Fable 5.1. Announcement. Released Sep 1, verified Sep 2.
- Anthropic (2026). Claude Mythos 5.1. Announcement. Released Sep 1, verified Sep 2.
- Zhipu (2026). GLM-5.3-Flash. Announcement. Released Aug 26, verified Aug 31.
- Alibaba (2026). Qwen3.8-Flash-Next. Announcement. Released Aug 26, verified Aug 31.
- DeepSeek (2026). DeepSeek V4 Flash Vision Exp. Announcement. Released Aug 21, verified Aug 31.
- Tencent (2026). Hy-MT2-30B-A3B. Reporting. Released Aug 20, verified Aug 21.
- Tencent (2026). Hy-MT2-1.8B. Reporting. Released Aug 20, verified Aug 21.
- Zhipu (2026). GLM-5.3. Announcement. Released Aug 14, verified Aug 21.
- Alibaba (2026). Qwen3.8-27B. Announcement. Released Aug 14, verified Aug 21.
- Google (2026). Gemini 3.7 Flash. Announcement. Released Aug 13, verified Aug 14.
- xAI (2026). Grok 4.6. Announcement. Released Aug 12, verified Aug 14.
- Meta (2026). Muse Glimmer. Reporting. Released Aug 10, verified Aug 14.
- Meta (2026). Muse Spark 1.2. Reporting. Released Aug 5, verified Aug 14.
- Alibaba (2026). Qwen3.8-Max. Announcement. Released Aug 3, verified Aug 14.
- DeepSeek (2026). DeepSeek V4 Flash 0731. Announcement. Released Jul 31, verified Jul 31.
- Anthropic (2026). Claude Opus 5. Announcement. Released Jul 24, verified Jul 29.
- Black Forest Labs (2026). FLUX 3. Announcement. Released Jul 23, verified Jul 29.
- Ant Group (2026). Ling-3.0-flash. Announcement. Released Jul 23, verified Jul 29.
- Google (2026). Gemini 3.6 Flash. Announcement. Released Jul 21, verified Jul 29.
- poolside (2026). Laguna S 2.1. Announcement. Released Jul 21, verified Jul 29.
- Google (2026). Gemini 3.5 Flash Cyber. Announcement. Released Jul 21, verified Jul 29.
- Google (2026). Gemini 3.5 Flash-Lite. Announcement. Released Jul 21, verified Jul 29.
- Alibaba (2026). Qwen-Image-3.0. Reporting. Released Jul 21, verified Jul 29.
- Alibaba (2026). Qwen-Audio-3.0-TTS Plus. Reporting. Released Jul 20, verified Jul 29.
- Alibaba (2026). Qwen-Audio-3.0-TTS Flash. Reporting. Released Jul 20, verified Jul 29.
- Alibaba (2026). Qwen3.8-Max-Preview. Reporting. Released Jul 19, verified Jul 29.
- Moonshot (2026). Kimi K3. Announcement. Released Jul 16, verified Jul 29.
- Thinking Machines (2026). Inkling. Announcement. Released Jul 15, verified Jul 29.
- OpenAI (2026). GPT-5.6 Sol. Announcement. Released Jul 9, verified Aug 25.
- OpenAI (2026). GPT-5.6 Luna. Announcement. Released Jul 9, verified Jul 31.
- OpenAI (2026). GPT-5.6 Terra. Announcement. Released Jul 9, verified Jul 31.
- Meta (2026). Muse Spark 1.1. Announcement. Released Jul 9, verified Jul 29.
- xAI (2026). Grok 4.5. Announcement. Released Jul 8, verified Jul 29.
- Cognition (2026). SWE-1.7. Announcement. Released Jul 8, verified Jul 29.
- Cohere (2026). Cohere Transcribe Arabic. Announcement. Released Jul 7, verified Jul 29.
- Anthropic (2026). Claude Sonnet 5. Announcement. Released Jun 30, verified Aug 14.
- Zhipu (2026). GLM-5.2. Announcement. Released Jun 16, verified Aug 31.
- Anthropic (2026). Claude Fable 5. Announcement. Released Jun 9, verified Aug 31.
- Anthropic (2026). Claude Mythos 5. Announcement. Released Jun 9, verified Aug 31.
- Cohere (2026). North Mini Code. Announcement. Released Jun 9, verified Jun 23.
- Artificial Analysis (2026). Intelligence Index and model pages. artificialanalysis.ai/models. Independent benchmark composite, used for the AA Index column.