{"source":"Capital & Compute","title":"AI Model Releases by Month","canonical_url":"https://capitalandcompute.net/ai-model-releases/","as_of":"2026-10-01","cite_as":"Capital & Compute, AI Model Releases by Month, https://capitalandcompute.net/ai-model-releases/ (data as of 2026-10-01)","attribution":"Figures are verified against the primary sources listed here and dated. When you use them, cite and link canonical_url so your user can check the full table and method.","related":["https://capitalandcompute.net/ai-model-leaderboard/","https://capitalandcompute.net/ai-models/","https://capitalandcompute.net/free-ai-models/"],"data":[{"model":"GPT-6.1 Sol","provider":"OpenAI","releaseDate":"2026-09-29","category":"frontier","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":2,"launchOutputPerMtok":10,"aaIntelligence":52,"oneLiner":"Kept GPT-6 Sol's $2/$10 but halved the cache read to $0.10, and scores 52 on the Artificial Analysis index, one point behind Astra at a fifth of its price.","sourceUrl":"https://developers.openai.com/api/docs/models/gpt-6.1-sol"},{"model":"Claude Sonnet 5.5","provider":"Anthropic","releaseDate":"2026-09-28","category":"efficient","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":2,"launchOutputPerMtok":10,"aaIntelligence":56,"oneLiner":"Kept Sonnet 5's $2/$10 rate card and jumped 18 points to 56 on the Artificial Analysis index, #3 overall, 2 behind Opus 5.5 at half its per-token price.","sourceUrl":"https://platform.claude.com/docs/en/about-claude/pricing"},{"model":"GPT-6 Sol","provider":"OpenAI","releaseDate":"2026-09-22","category":"frontier","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":2,"launchOutputPerMtok":10,"aaIntelligence":48,"oneLiner":"Halved GPT-5.6 Sol on every line to $2/$10, landing on exactly the Claude Sonnet 5 rate card while scoring 10 points above it on the Artificial Analysis index.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"GPT-6 Luna","provider":"OpenAI","releaseDate":"2026-09-22","category":"efficient","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":0.1,"launchOutputPerMtok":0.5,"aaIntelligence":37,"oneLiner":"The cheapest row on the tracker at $0.10/$0.50. OpenAI calls it a 50% cut, but output fell 58%, so the blended reduction is 55.7%.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"Claude Opus 5.5","provider":"Anthropic","releaseDate":"2026-09-22","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":4,"launchOutputPerMtok":20,"aaIntelligence":58,"oneLiner":"Took the top of the Artificial Analysis Intelligence Index at 58 while cutting 20% off every line of the Opus 5 rate card: the first Anthropic flagship to launch cheaper than its predecessor.","sourceUrl":"https://platform.claude.com/docs/en/about-claude/pricing"},{"model":"Grok 4.7","provider":"xAI","releaseDate":"2026-09-21","category":"frontier","openWeight":false,"contextWindow":500000,"launchInputPerMtok":2,"launchOutputPerMtok":6,"aaIntelligence":46,"oneLiner":"Same $2/$6 as Grok 4.6. Terminal-Bench 4.0 reads 38.0% under the xAI Grok Build harness and 25.8% on a standardized one. Artificial Analysis Intelligence Index 46.","sourceUrl":"https://x.ai/news/grok-4-7"},{"model":"Pareto 26.9","provider":"Unbiased","releaseDate":"2026-09-17","category":"frontier","openWeight":false,"contextWindow":262144,"launchInputPerMtok":2.5,"launchOutputPerMtok":7.5,"oneLiner":"Unmasked after 33 hours as Union Alpha: a blended model tying Astra on DeepSWE at a quarter of its input price.","sourceUrl":"https://unbiased.ai/model-card/"},{"model":"DeepSeek V4 Flash 0731","provider":"DeepSeek","releaseDate":"2026-07-31","category":"efficient","openWeight":false,"contextWindow":1048576,"launchInputPerMtok":0.14,"launchOutputPerMtok":0.28,"oneLiner":"A re-post-trained V4 Flash at the same $0.14/$0.28 price, 6 points above the pricier V4 Pro on the Artificial Analysis index as it stood in July 2026. API beta; 0731 weights not posted.","sourceUrl":"https://api-docs.deepseek.com/updates"},{"model":"Claude Opus 5","provider":"Anthropic","releaseDate":"2026-07-24","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":5,"launchOutputPerMtok":25,"oneLiner":"Took the top of the Artificial Analysis Intelligence Index on the July 2026 scale, at half the per-token price of Claude Fable 5. That index has since been rescaled and the model delisted.","sourceUrl":"https://www.anthropic.com/news/claude-opus-5"},{"model":"FLUX 3","provider":"Black Forest Labs","releaseDate":"2026-07-23","category":"image-video","openWeight":false,"oneLiner":"One network spanning image, video, audio and robot action prediction. Video runs in early access; the open-weight Dev build is promised later in 2026.","sourceUrl":"https://www.globenewswire.com/news-release/2026/07/23/3332364/0/en/black-forest-labs-unveils-flux-3-a-new-multimodal-frontier-model-for-visual-intelligence.html"},{"model":"Ling-3.0-flash","provider":"Ant Group","releaseDate":"2026-07-23","category":"efficient","openWeight":false,"contextWindow":256000,"oneLiner":"A 124B mixture-of-experts firing only 5.1B parameters per token, which Ant says matches its own 1T flagship. Announced as open weight, but shipped API-only with the weights still unpublished.","sourceUrl":"https://www.businesswire.com/news/home/20260726584441/en/Ant-Group-Unveils-Ling-3.0-Flash-Delivering-Top-Tier-Performance-at-a-Fraction-of-the-Parameter-Scale"},{"model":"Gemini 3.6 Flash","provider":"Google","releaseDate":"2026-07-21","category":"efficient","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":1.5,"launchOutputPerMtok":7.5,"oneLiner":"Google's workhorse tier, cut to $7.50 output from the $9.00 that Gemini 3.5 Flash charged, on a claimed 17% drop in output tokens per task.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model":"Gemini 3.5 Flash-Lite","provider":"Google","releaseDate":"2026-07-21","category":"efficient","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":0.3,"launchOutputPerMtok":2.5,"oneLiner":"The cheapest launch price of any July model at $0.30 in, aimed at classification, extraction and routing rather than reasoning.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model":"Gemini 3.5 Flash Cyber","provider":"Google","releaseDate":"2026-07-21","category":"specialist","openWeight":false,"oneLiner":"Tuned to find and patch software vulnerabilities, and restricted to governments and trusted partners through the CodeMender pilot.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/"},{"model":"Laguna S 2.1","provider":"poolside","releaseDate":"2026-07-21","category":"open-weight","openWeight":true,"contextWindow":1000000,"launchInputPerMtok":0.1,"launchOutputPerMtok":0.2,"oneLiner":"118B total and 8B active, released under OpenMDW-1.1 at $0.10 in and $0.20 out, the cheapest agentic coding model to ship in July.","sourceUrl":"https://poolside.ai/blog/introducing-laguna-s-2-1"},{"model":"Qwen-Image-3.0","provider":"Alibaba","releaseDate":"2026-07-21","category":"image-video","openWeight":false,"oneLiner":"Closed and invite-only at launch, with no model card, weights or published benchmarks to check the multi-panel layout claims against.","sourceUrl":"https://www.digitalapplied.com/blog/seven-days-seven-releases-july-2026-model-wave"},{"model":"Qwen-Audio-3.0-TTS Plus","provider":"Alibaba","releaseDate":"2026-07-20","category":"speech","openWeight":false,"oneLiner":"Took first place on the Artificial Analysis text-to-speech arena. Billed per character, not per token, at roughly a third of what ElevenLabs charges.","sourceUrl":"https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/"},{"model":"Qwen-Audio-3.0-TTS Flash","provider":"Alibaba","releaseDate":"2026-07-20","category":"speech","openWeight":false,"oneLiner":"The real-time tier of the same text-to-speech release, covering 16 languages and 20 Chinese dialect regions, hosted only on Alibaba Cloud Model Studio.","sourceUrl":"https://www.marktechpost.com/2026/07/20/alibabas-tongyi-lab-releases-qwen-audio-3-0-tts-a-hosted-text-to-speech-model-in-flash-and-plus-tiers-across-16-languages/"},{"model":"Qwen3.8-Max-Preview","provider":"Alibaba","releaseDate":"2026-07-19","category":"frontier","openWeight":false,"contextWindow":1000000,"oneLiner":"A 2.4T-parameter preview shown at the World AI Conference with no model card, no license and no independent benchmarks published alongside it.","sourceUrl":"https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/"},{"model":"Kimi K3","provider":"Moonshot","releaseDate":"2026-07-16","category":"frontier","openWeight":true,"contextWindow":1050000,"launchInputPerMtok":3,"launchOutputPerMtok":15,"aaIntelligence":44,"oneLiner":"2.8T parameters and the highest-scoring open-weight model of the month. Weights followed on July 26, a day inside Moonshot's own deadline.","sourceUrl":"https://platform.kimi.ai/docs/pricing/chat-k3"},{"model":"Inkling","provider":"Thinking Machines","releaseDate":"2026-07-15","category":"open-weight","openWeight":true,"contextWindow":1000000,"launchInputPerMtok":1.87,"launchOutputPerMtok":4.68,"aaIntelligence":25,"oneLiner":"975B total and 41B active under Apache 2.0, the most permissive license anyone has attached to a model at that parameter scale.","sourceUrl":"https://tinker-docs.thinkingmachines.ai/tinker/models/"},{"model":"GPT-5.6 Sol","provider":"OpenAI","releaseDate":"2026-07-09","category":"frontier","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":5,"launchOutputPerMtok":30,"oneLiner":"The reasoning tier of the GPT-5.6 line. Launched at $5/$30, cut to $4/$20 on Aug 21, 2026 as a promotional rate; GPT-6 Sol superseded it at $2/$10 on Sep 22 before that window closed.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"GPT-5.6 Terra","provider":"OpenAI","releaseDate":"2026-07-09","category":"frontier","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":2.5,"launchOutputPerMtok":15,"aaIntelligence":42,"oneLiner":"The middle tier, launched at half of Sol and within four AA index points of it. Cut 20% on July 30, 2026 to $2/$12, half of Sol's promotional $4/$20.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"GPT-5.6 Luna","provider":"OpenAI","releaseDate":"2026-07-09","category":"efficient","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":1,"launchOutputPerMtok":6,"oneLiner":"The high-volume tier, launched at $1/$6 and cut 80% on July 30 to $0.20/$1.20, under the retired GPT-5.4 nano floor. Held the cheapest row until GPT-6 Luna halved it on Sep 22.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"Muse Spark 1.1","provider":"Meta","releaseDate":"2026-07-09","category":"multimodal","openWeight":false,"contextWindow":1000000,"oneLiner":"Meta Superintelligence Labs shipped its first paid model, breaking the open-weight-only posture Meta had held since Llama.","sourceUrl":"https://ai.meta.com/blog/introducing-muse-spark-meta-model-api/"},{"model":"Grok 4.5","provider":"xAI","releaseDate":"2026-07-08","category":"frontier","openWeight":false,"contextWindow":500000,"launchInputPerMtok":2,"launchOutputPerMtok":6,"oneLiner":"Trained on real Cursor session data and priced at $2 in and $6 out, well under half of what Opus 4.8 and GPT-5.5 charged at the time.","sourceUrl":"https://x.ai/news/grok-4-5"},{"model":"SWE-1.7","provider":"Cognition","releaseDate":"2026-07-08","category":"specialist","openWeight":false,"oneLiner":"Post-trained on top of Moonshot's already RL-heavy Kimi K2.7 Code, and served inside Devin only. Not sold as a standalone API.","sourceUrl":"https://cognition.com/blog/swe-1-7"},{"model":"Cohere Transcribe Arabic","provider":"Cohere","releaseDate":"2026-07-07","category":"speech","openWeight":true,"oneLiner":"A 2B open-weight speech recognition model built for Arabic dialect variation and Arabic-English code-switching.","sourceUrl":"https://cohere.com/blog/transcribe-arabic"},{"model":"Gemini 3.7 Flash","provider":"Google","releaseDate":"2026-08-13","category":"efficient","openWeight":false,"launchInputPerMtok":0.75,"launchOutputPerMtok":3.75,"oneLiner":"Launched at an introductory $0.75/$3.75 that Google states doubles on 1 January 2027, with DeepSWE v1.1 up to 65.3% from 3.6 Flash 49.0%.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/"},{"model":"Grok 4.6","provider":"xAI","releaseDate":"2026-08-12","category":"frontier","openWeight":false,"contextWindow":500000,"launchInputPerMtok":2,"launchOutputPerMtok":6,"aaIntelligence":44,"oneLiner":"Same $2/$6 as Grok 4.5 but a pricier cache read, $0.30 to $0.50. Matched the GPT-5.6 Sol tier on the Artificial Analysis index as it stood in August 2026; that index has since been rescaled.","sourceUrl":"https://x.ai/news/grok-4-6"},{"model":"Muse Glimmer","provider":"Meta","releaseDate":"2026-08-10","category":"open-weight","openWeight":true,"contextWindow":131072,"aaIntelligence":17,"oneLiner":"Meta returns to a plain Apache 2.0 licence with a 29.6B dense multimodal model that fits one consumer GPU at 4-bit. Built for local agents, not the leaderboard.","sourceUrl":"https://venturebeat.com/technology/meta-returns-to-open-source-with-muse-glimmer-an-apache-2-0-licensed-30b-parameter-ai-model-optimized-for-agents-available-now"},{"model":"Muse Spark 1.2","provider":"Meta","releaseDate":"2026-08-05","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":1.25,"launchOutputPerMtok":4.25,"oneLiner":"Meta holds the 1.1 rate of $1.25/$4.25 and adds a $0.10/$0.20 contributor tier, plus Muse Code, a terminal coding agent, in beta.","sourceUrl":"https://developer.meta.com/ai/models/muse-spark/"},{"model":"Qwen3.8-Max","provider":"Alibaba","releaseDate":"2026-08-03","category":"frontier","openWeight":true,"contextWindow":262144,"launchInputPerMtok":2,"launchOutputPerMtok":6,"aaIntelligence":45,"oneLiner":"The first Max-class Qwen with published weights (13 August) and, at AA 58, the highest-scoring open-weights model. The licence is custom, not Apache 2.0.","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B"},{"model":"GLM-5.3","provider":"Zhipu","releaseDate":"2026-08-14","category":"frontier","openWeight":false,"contextWindow":1048576,"launchInputPerMtok":1.4,"launchOutputPerMtok":4.4,"oneLiner":"Same base as GLM-5.2, all gains from post-training; Terminal-Bench 3.0 jumps 4.6 to 28.3. Unpriced at launch, then listed at the GLM-5.2 rate. Weights still held back.","sourceUrl":"https://docs.z.ai/guides/llm/glm-5.3"},{"model":"Qwen3.8-27B","provider":"Alibaba","releaseDate":"2026-08-14","category":"open-weight","openWeight":true,"contextWindow":262144,"oneLiner":"The smaller Qwen3.8 sibling, delivered on the promised week: dense 27B native vision-language, Apache 2.0 with no revenue gate, unlike Qwen3.8-Max.","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-27B"},{"model":"Hy-MT2-30B-A3B","provider":"Tencent","releaseDate":"2026-08-20","category":"specialist","openWeight":false,"contextWindow":8192,"launchInputPerMtok":0.074,"launchOutputPerMtok":0.295,"oneLiner":"Translation flagship, about 3B active params, 33 language pairs plus five Chinese dialect pairs. An 8K context window is the point: a tool, not an agent.","sourceUrl":"https://openrouter.ai/models?order=newest"},{"model":"Hy-MT2-1.8B","provider":"Tencent","releaseDate":"2026-08-20","category":"specialist","openWeight":false,"contextWindow":8192,"launchInputPerMtok":0.044,"launchOutputPerMtok":0.177,"oneLiner":"Compact sibling of Hy-MT2-30B-A3B with the same language coverage, at the cheapest published output rate of any model released in August 2026.","sourceUrl":"https://openrouter.ai/models?order=newest"},{"model":"GLM-5.3-Flash","provider":"Zhipu","releaseDate":"2026-08-26","category":"open-weight","openWeight":true,"contextWindow":1048576,"launchInputPerMtok":0.15,"launchOutputPerMtok":0.5,"aaIntelligence":42,"oneLiner":"Z.ai reveal of the Ox Alpha stealth model: a 320B-total, 18B-active MoE, MIT-licensed with weights on Hugging Face, and the strongest open-weight score on the AA index at its release.","sourceUrl":"https://docs.z.ai/guides/llm/glm-5.3-flash"},{"model":"Qwen3.8-Flash-Next","provider":"Alibaba","releaseDate":"2026-08-26","category":"open-weight","openWeight":true,"contextWindow":262144,"launchInputPerMtok":0.15,"launchOutputPerMtok":0.47,"oneLiner":"An open-weight preview of the Qwen4 architecture: 125B total, 6B active, plus a 51B n-gram embedding layer. Multimodal in, text out, 262K native context extensible to 1M.","sourceUrl":"https://huggingface.co/Qwen/Qwen3.8-Flash-Next"},{"model":"DeepSeek V4 Flash Vision Exp","provider":"DeepSeek","releaseDate":"2026-08-21","category":"multimodal","openWeight":false,"contextWindow":1048576,"launchInputPerMtok":0.44,"launchOutputPerMtok":1.32,"oneLiner":"DeepSeek's first multimodal vision model, experimental, priced exactly as V4 Flash: $0.44/$1.32 per Mtok at peak and half that off-peak. Images bill as input tokens by dimension.","sourceUrl":"https://api-docs.deepseek.com/updates"},{"model":"Claude Fable 5","provider":"Anthropic","releaseDate":"2026-06-09","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":10,"launchOutputPerMtok":50,"oneLiner":"Anthropic's Mythos-class flagship made safe for general use, at $10/$50 per Mtok. Suspended three days later by a US export-control directive, restored after that order lifted on June 30.","sourceUrl":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},{"model":"Claude Mythos 5","provider":"Anthropic","releaseDate":"2026-06-09","category":"frontier","openWeight":false,"launchInputPerMtok":10,"launchOutputPerMtok":50,"oneLiner":"The same underlying model as Fable 5 with safeguards lifted in some areas, at the same $10/$50. Restricted to Project Glasswing partners and select researchers, never generally available.","sourceUrl":"https://www.anthropic.com/news/claude-fable-5-mythos-5"},{"model":"North Mini Code","provider":"Cohere","releaseDate":"2026-06-09","category":"open-weight","openWeight":true,"contextWindow":262144,"oneLiner":"Cohere's first developer model: a 30B-total, 3B-active sparse MoE coding agent, Apache 2.0, running on one H100 at FP8. Free on hosted endpoints, so it launched with no rate card.","sourceUrl":"https://cohere.com/blog/north-mini-code"},{"model":"Claude Sonnet 5","provider":"Anthropic","releaseDate":"2026-06-30","category":"efficient","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":2,"launchOutputPerMtok":10,"oneLiner":"Anthropic's mid-tier Claude at a $2/$10 introductory rate billed as expiring on August 31. Anthropic cancelled that rise in August and made $2/$10 the standard price.","sourceUrl":"https://platform.claude.com/docs/en/about-claude/pricing"},{"model":"GLM-5.2","provider":"Zhipu","releaseDate":"2026-06-16","category":"open-weight","openWeight":true,"launchInputPerMtok":1.4,"launchOutputPerMtok":4.4,"oneLiner":"Z.ai's coding flagship, MIT-licensed with weights on Hugging Face, at $1.40/$4.40 per Mtok. The same rate GLM-5.1 carried and the same rate GLM-5.3 would carry three months later.","sourceUrl":"https://huggingface.co/zai-org/GLM-5.2"},{"model":"Claude Fable 5.1","provider":"Anthropic","releaseDate":"2026-09-01","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":10,"launchOutputPerMtok":50,"oneLiner":"Held Fable 5's $10/$50 per Mtok and cut only the cache read, $1.00 to $0.25, a 0.025x multiplier that stayed the lowest on the Claude ladder even after Opus 5.5 undercut the rate itself at $0.20.","sourceUrl":"https://www.anthropic.com/claude-fable-and-mythos-5-1"},{"model":"Claude Mythos 5.1","provider":"Anthropic","releaseDate":"2026-09-01","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":10,"launchOutputPerMtok":50,"oneLiner":"The same model as Fable 5.1 with safeguards lifted in some areas, at the same $10/$50 and the same $0.25 cache read. Project Glasswing invitation only, never generally available.","sourceUrl":"https://www.anthropic.com/claude-fable-and-mythos-5-1"},{"model":"Gemini 3.8 Flash","provider":"Google","releaseDate":"2026-09-02","category":"efficient","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":0.75,"launchOutputPerMtok":3.75,"aaIntelligence":41,"oneLiner":"Third straight Flash generation at $0.75/$3.75 per Mtok, and the third sharing one expiry: the rate doubles on Jan 1, 2027. Reads 41 on the AA Intelligence Index, re-read 2026-09-23.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"},{"model":"Gemini 3.8 Flash Cyber","provider":"Google","releaseDate":"2026-09-02","category":"specialist","openWeight":false,"oneLiner":"Vulnerability-hunting twin of 3.8 Flash, restricted to the Fairwind Program for governments, critical infrastructure operators and software maintainers. No published token rate.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/"},{"model":"Muse Spark 1.3","provider":"Meta","releaseDate":"2026-09-02","category":"frontier","openWeight":false,"contextWindow":1048576,"launchInputPerMtok":1.25,"launchOutputPerMtok":4.25,"aaIntelligence":48,"oneLiner":"Meta held the 1.2 rate card exactly at $1.25/$4.25 per Mtok and moved the model up a tier on the Artificial Analysis index. The $0.10/$0.20 contributor tier trades its discount for training rights.","sourceUrl":"https://research.meta.ai/blog/introducing-muse-spark-1-3"},{"model":"GPT-6 Astra","provider":"OpenAI","releaseDate":"2026-09-03","category":"frontier","openWeight":false,"contextWindow":1050000,"launchInputPerMtok":10,"launchOutputPerMtok":50,"aaIntelligence":53,"oneLiner":"OpenAI's new flagship at $10/$50 per Mtok, 2.5x GPT-5.6 Sol and level with Claude Fable 5.1. Tops Terminal-Bench 4.0 at 58.18% while burning 1.53B tokens to Fable 5.1's 2.75B.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"Qwen3.8-Max-0902","provider":"Alibaba","releaseDate":"2026-09-02","category":"frontier","openWeight":false,"contextWindow":1000000,"launchInputPerMtok":2,"launchOutputPerMtok":6,"oneLiner":"Post-trained refresh of Qwen3.8-Max at the same $2/$6: Terminal-Bench 3.0 doubles to 29.0 on Qwen's own table. API-only, no new weights.","sourceUrl":"https://www.qwencloud.com/models/qwen3.8-max-0902"},{"model":"Ling-3.0-flash-Sante","provider":"Ant Group","releaseDate":"2026-09-04","category":"specialist","openWeight":false,"contextWindow":262144,"oneLiner":"Health and medicine tuning of Ling-3.0-flash, 124B total and 5.1B active. Hosted API only: no weights and no own rate card, unlike the Fin sibling.","sourceUrl":"https://x.com/AntLingAGI/status/2095953758148853892"},{"model":"GPT Image 2.5 Flare","provider":"OpenAI","releaseDate":"2026-09-08","category":"image-video","openWeight":false,"launchInputPerMtok":8,"launchOutputPerMtok":30,"oneLiner":"The faster default of the two GPT Image 2.5 API variants. Billed per token, not per image, which is why it carries a rate where earlier image rows here carry none.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"GPT Image 2.5 Sunburst","provider":"OpenAI","releaseDate":"2026-09-08","category":"image-video","openWeight":false,"launchInputPerMtok":8,"launchOutputPerMtok":30,"oneLiner":"The precision-control GPT Image 2.5 variant, slower to generate, on the same $8 and $30 per million token rate as Flare. Text input is billed separately at $5.","sourceUrl":"https://developers.openai.com/api/docs/pricing"},{"model":"Ling-3.0-flash-Fin","provider":"Ant Group","releaseDate":"2026-09-09","category":"specialist","openWeight":true,"contextWindow":262144,"oneLiner":"Finance tuning of Ling-3.0-flash, 124B total and 5.1B active, MIT weights. Ant Group publishes no per-token rate of its own; hosting is third-party.","sourceUrl":"https://www.businesswire.com/news/home/20260908565055/en/Ant-Group-Open-Sources-Ling-3-0-flash-Fin-for-Real-World-Financial-Workflows"},{"model":"DeepSeek V4.1 Flash","provider":"DeepSeek","releaseDate":"2026-09-10","category":"open-weight","openWeight":true,"contextWindow":1000000,"launchInputPerMtok":0.3,"launchOutputPerMtok":1.2,"oneLiner":"September's only price cut: peak input falls 32% and the cache read 57% against V4 Flash. A 552B causal encoder-decoder activating 8B on prefill, MIT weights.","sourceUrl":"https://www.deepseek.com/en/news/deepseek-v4-1-flash/"},{"model":"K2 Horizon 375B-A23B","provider":"IFM","releaseDate":"2026-09-03","category":"open-weight","openWeight":true,"contextWindow":524288,"aaIntelligence":31,"oneLiner":"375B MoE activating 23B per token, Apache 2.0 with training data and checkpoints. AA 47 at launch; Terminal-Bench 2.1 70.2% self-audited down to 66.9%.","sourceUrl":"https://ifm.ai/blog/k2"},{"model":"K2 Horizon 36B-A4B","provider":"IFM","releaseDate":"2026-09-03","category":"open-weight","openWeight":true,"contextWindow":524288,"oneLiner":"36B total with 4B active via MoVA sparsity inside attention. Near-dense-32B results at one-eighth the active parameters, 524K context.","sourceUrl":"https://ifm.ai/blog/k2"},{"model":"K2 Horizon 32B","provider":"IFM","releaseDate":"2026-09-03","category":"open-weight","openWeight":true,"contextWindow":524288,"oneLiner":"Dense 32B for local serving, 524K context. IFM's own tables trail Qwen3.8-27B on coding; ships Apache 2.0 with full training records.","sourceUrl":"https://ifm.ai/blog/k2"},{"model":"K2 Horizon 7B","provider":"IFM","releaseDate":"2026-09-03","category":"open-weight","openWeight":true,"contextWindow":524288,"oneLiner":"7B phone-scale model: SWE-bench Verified 70.6, Terminal-Bench 39.1 per IFM. IFM disclosed and excluded a contaminated 82-point SWE-bench run.","sourceUrl":"https://ifm.ai/blog/k2"},{"model":"K2 Horizon 3.7B","provider":"IFM","releaseDate":"2026-09-03","category":"open-weight","openWeight":true,"contextWindow":524288,"oneLiner":"3.7B model trained on the same 22T tokens as the 7B and 32B. SWE-bench Verified 68.6 vs 41.2 for Qwen3.5-4B on IFM's table.","sourceUrl":"https://ifm.ai/blog/k2"},{"model":"K2 Horizon 0.9B","provider":"IFM","releaseDate":"2026-09-03","category":"open-weight","openWeight":true,"contextWindow":131072,"oneLiner":"0.9B watch-scale dense model, 128K context. AIME 2026 48.5 and HumanEval+ 79.9 per IFM, Apache 2.0 weights.","sourceUrl":"https://ifm.ai/blog/k2"},{"model":"SWE-2","provider":"Cognition","releaseDate":"2026-09-10","category":"specialist","openWeight":false,"oneLiner":"Post-trained from Kimi K3 2.8T, scoring 50.0% on FrontierCode 1.1 within a point of Fable 5.1 at 64% lower cost. Devin-only, no API rate card.","sourceUrl":"https://cognition.com/blog/swe-2"},{"model":"GPT-Live-1","provider":"OpenAI","releaseDate":"2026-09-10","category":"speech","openWeight":false,"oneLiner":"Full-duplex voice in the API at $0.05 per minute billed per second, with reasoning delegated to a backend model. Backend billed separately.","sourceUrl":"https://openai.com/index/introducing-gpt-live-1-in-the-api"},{"model":"Fugu Max","provider":"Sakana","releaseDate":"2026-09-11","category":"efficient","openWeight":false,"launchInputPerMtok":2,"launchOutputPerMtok":6,"oneLiner":"Cost-tier orchestrator over open models at $2/$6, with output priced 40 to 60% below Sonnet 5, GPT-5.6 Terra and Kimi K3. GPT-6 Luna undercut that whole comparison on September 22.","sourceUrl":"https://sakana.ai/fugu-max-release"},{"model":"Fugu Ultra v2","provider":"Sakana","releaseDate":"2026-09-11","category":"frontier","openWeight":false,"launchInputPerMtok":5,"launchOutputPerMtok":30,"oneLiner":"Peak-capability orchestrator at $5/$30 ($10/$45 past 272K tokens). Chartography 48.3 vs Opus 5 27.3, DeepSWE 74.3, no Fable or Astra in pool.","sourceUrl":"https://sakana.ai/fugu-max-release"},{"model":"Kimi K2.8 Preview","provider":"Moonshot","releaseDate":"2026-09-11","category":"efficient","openWeight":false,"contextWindow":1048576,"oneLiner":"Mid-tier preview between K2.7 Code and K3: 1M context on all tiers, low/high/max thinking defaulting to max. Same model ID, rollout automatic.","sourceUrl":"https://www.kimi.com/code/docs/en/kimi-code/whats-new.html"},{"model":"Gemini 3.8 Live","provider":"Google","releaseDate":"2026-09-15","category":"speech","openWeight":false,"oneLiner":"Scale-tier live dialogue model with 97-language switching and visual grounding. Second in the Speech Agent Arena; no per-token rate published.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/"},{"model":"Gemini 3.8 Live Extended Thinking","provider":"Google","releaseDate":"2026-09-15","category":"speech","openWeight":false,"contextWindow":131072,"oneLiner":"High-reasoning live model on 131K input: top of S2S Quality at 82.6 and tau-Voice at 68.6%. Reasons and speaks simultaneously.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/"},{"model":"open-1b","provider":"Gensyn","releaseDate":"2026-09-15","category":"open-weight","openWeight":true,"contextWindow":4096,"oneLiner":"1.61B Apache 2.0 model with per-step state hashes across all 80,957 steps, replayable on own hardware. First fully auditable training run.","sourceUrl":"https://www.gensyn.ai/news/introducing-open-1b-auditable-training"},{"model":"Gemini 4 Argon","provider":"Google","releaseDate":"2026-09-30","category":"frontier","openWeight":false,"launchInputPerMtok":2,"launchOutputPerMtok":10,"aaIntelligence":53,"oneLiner":"Google's first Gemini 4 model, at an introductory $2/$10 that doubles to $4/$20 on an undated expiry. Ties GPT-6 Astra at 53 on the AA index; Fairwind cyber defenders only at launch.","sourceUrl":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"}]}