New AI Models Released in August 2026: The Full List
Every new AI model released in August 2026 with dates, verified per-token prices and primary sources, plus the price changes that landed alongside them.
Twelve distinct model releases landed in August 2026: Qwen3.8-Max from Alibaba (August 3), Meta’s Muse Spark 1.2 (August 5) and Muse Glimmer (August 10), xAI’s Grok 4.6 (August 12), Google’s Gemini 3.7 Flash (August 13), Z.ai’s GLM-5.3 and Alibaba’s open-weight Qwen3.8-27B (both August 14), Tencent’s two Hy-MT2 translation models (August 20), DeepSeek’s experimental V4 Flash Vision (August 21), and then a closing pair on August 26: Z.ai’s GLM-5.3-Flash and Alibaba’s Qwen3.8-Flash-Next. Every date and rate below is verified against the lab’s own page, and the price is stated as the provider states it, including when that price has an expiry date attached.
The month’s real story is not in the release list. It is that three separate labs changed what a published price means. Google’s newest Flash tier carries a rate that doubles on 1 January 2027. Anthropic cancelled a scheduled increase and made a discount permanent. DeepSeek is replacing a flat rate with peak and off-peak billing on 16 August, and the new off-peak rate is more than double the price it charges today. Prices moved in both directions in the same fortnight. This page is the August installment of a monthly series; the complete dated list lives in the record of AI model releases by month, and current per-token rates and upcoming models are in the AI model tracker.
Every AI model released in August 2026
| Model | Maker | Released | Price (in/out per Mtok) | What it is |
|---|---|---|---|---|
| Qwen3.8-Max | Alibaba | Aug 3 | $2 / $6 | 2.4T MoE flagship, open weights from Aug 13, custom licence |
| Muse Spark 1.2 | Meta | Aug 5 | $1.25 / $4.25 | Paid agentic model, 1M context, plus Muse Code terminal agent |
| Muse Glimmer | Meta | Aug 10 | Not published | 29.6B dense multimodal, Apache 2.0, runs on one consumer GPU |
| Grok 4.6 | xAI | Aug 12 | $2 / $6 | Frontier model, 500K context, cached input raised to $0.50 |
| Gemini 3.7 Flash | Aug 13 | $0.75 / $3.75 | Efficient tier, introductory rate that doubles on Jan 1, 2027 | |
| GLM-5.3 | Z.ai | Aug 14 | $1.40 / $4.40 | Coding and agent flagship on the GLM-5.2 base, at the 5.2 rate |
| Qwen3.8-27B | Alibaba | Aug 14 | Weights only | 27B dense vision-language, Apache 2.0, 262K native context |
| Hy-MT2-30B-A3B | Tencent | Aug 20 | $0.074 / $0.295 | Translation flagship, 33 language pairs, 8K context, 3B active |
| Hy-MT2-1.8B | Tencent | Aug 20 | $0.044 / $0.177 | Compact translation model, the same language coverage |
| V4 Flash Vision Exp | DeepSeek | Aug 21 | $0.44 / $1.32 peak | Experimental vision model, priced exactly as V4 Flash |
| GLM-5.3-Flash | Z.ai | Aug 26 | $0.15 / $0.50 | The Ox Alpha reveal, 320B/18B multimodal MoE, MIT weights |
| Qwen3.8-Flash-Next | Alibaba | Aug 26 | $0.15 / $0.47 | Open-weight Qwen4 architecture preview, 125B total, 6B active |
Seven labs, and the shape of the month is three clusters. July’s pattern was three labs shipping into a single 48-hour window. August front-loaded instead: five frontier-tier releases between August 10 and August 14, two of them on the 14th, then a quiet stretch broken only by Tencent’s translation pair on the 20th and DeepSeek’s experimental vision model on the 21st, then a second frontier cluster on August 26 that closed the month. What arrived in the middle gap was not a model at all. It was a rate card for a model that had already shipped.
| Day of August 2026 | Releases | Models |
|---|---|---|
| 3 | 1 | Qwen3.8-Max |
| 5 | 1 | Muse Spark 1.2 |
| 10 | 1 | Muse Glimmer |
| 12 | 1 | Grok 4.6 |
| 13 | 1 | Gemini 3.7 Flash |
| 14 | 2 | GLM-5.3, Qwen3.8-27B |
| 20 | 2 | Hy-MT2-30B-A3B, Hy-MT2-1.8B |
| 21 | 1 | DeepSeek V4 Flash Vision Exp |
| 26 | 2 | GLM-5.3-Flash, Qwen3.8-Flash-Next |
The month a published price stopped being a single number
Three labs changed the meaning of their own rate card inside two weeks, in three different ways, and only one of those changes was a cut.
DeepSeek is the sharpest. Its API documentation states that from 16:00 UTC on August 16, 2026 it adopts peak and off-peak pricing, with off-peak set at half the peak rate, peak hours running 01:00 to 04:00 and 06:00 to 10:00 UTC. Read as a discount scheme that sounds generous. Read against the price it charges today, it is a substantial increase. DeepSeek V4 Pro currently costs $0.435 input and $0.87 output per million tokens. After the change the off-peak rate is $0.66 and $1.98, and the peak rate is $1.32 and $3.96. The cheapest hour of the new schedule is 2.3x the current output price, and the most expensive is 4.6x.
The cheapest hour of DeepSeek's new schedule costs more than twice what the same model costs today.
Google went the other way, twice, with a clock attached. Gemini 3.7 Flash launched on August 13 at $0.75 input and $3.75 output per million tokens, and Google’s Gemini API pricing page states plainly that this is introductory: on January 1, 2027 it becomes $1.50 and $7.50, with the context-cache rate going from $0.075 to $0.15. The same page now lists Gemini 3.6 Flash, released in late July at $1.50/$7.50, at the identical $0.75/$3.75. That is a 50% cut roughly three weeks after launch, and it also expires at the end of the year.
Anthropic did the rarest thing: it cancelled a planned increase. Claude Sonnet 5 launched on June 30 at $2/$10 per million tokens, described at the time as introductory pricing through August 31, 2026, reverting to $3/$15 on September 1. Anthropic’s pricing page now carries a note stating that the $2/$10 rate “is now the standard price” and that “the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.” Anyone who planned a September migration off Sonnet 5 to avoid a 50% rise can stop.
| Item | Higher rate | Lower rate |
|---|---|---|
| DeepSeek V4 Flash (by hour) | $1.32 | $0.66 |
| DeepSeek V4 Pro (by hour) | $3.96 | $1.98 |
| Gemini 3.6 Flash (by date) | $7.50 | $3.75 |
| Gemini 3.7 Flash (by date) | $7.50 | $3.75 |
| Claude Sonnet 5 (cancelled) | $15.00 | $10.00 |
Qwen3.8-Max: the top open-weights score, and a licence that is not open source (August 3)
Alibaba announced Qwen3.8-Max on August 3 on the Model Studio API and published the weights on August 13 as Qwen3.8-2.4T-A95B on Hugging Face. It is a 2.4 trillion-parameter mixture of experts with 95 billion active, 512 experts with 11 activated per token, and a 262,144-token native context extensible to roughly 1.01 million. It is the first Max-class Qwen model to ship downloadable weights, and on the independent Artificial Analysis Intelligence Index it scores 58, the highest of any open-weights model on that board as of August 14.
Qwen’s own published benchmarks put it at 92.6 on GPQA Diamond, 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro, 56.6 on Deep-SWE 1.1 and 93.0 on PaperBench. Those are vendor-reported and have not been independently reproduced.
The licence is where the coverage has been sloppy. The model card states a custom qwen3.8-max licence, not Apache 2.0. Reading the licence text itself, two thresholds matter. Above 100 million monthly active users or US$20 million monthly revenue, the model name must be prominently displayed. More consequentially, if aggregate revenue exceeds US$50 million over any twelve consecutive months for a model-as-a-service or AI work-assistant business, a separate licence from Qwen is required before commercial use. Internal use is explicitly exempt, so a company running it on its own workloads is unaffected. A company reselling inference is not.
Muse Glimmer: Meta’s genuinely permissive release (August 10)
Meta published Muse Glimmer on Hugging Face under a plain apache-2.0 licence, with no acceptable-use addendum and no revenue threshold. It is a dense causal transformer of roughly 29.6 billion parameters including a perception encoder, with a 131,072-token context, text and image input, text output, more than 100 languages and a January 4, 2026 knowledge cutoff. It is distilled from Muse Spark and aimed at autonomous agentic work on consumer hardware: quantized to 4 bits it drops under 20 GB and fits a single 24 or 32 GB GPU.
Meta’s published benchmarks are strong: 94.7% on AIME 2026, 83.5% on GPQA Diamond, 76.0% on SWE-Bench Verified, 75.5% on MCP Atlas. Artificial Analysis scores it 35 on the Intelligence Index at high effort, which is the widest gap between vendor-reported and independent scoring of any August release, and worth weighing before adopting it on the strength of the launch table alone. There is no first-party per-token API rate, so no price appears in our registry.
Note also what did not happen: there was no Llama release in this window at all. Meta’s open-weights activity has moved to the Muse family, and coverage that still frames Meta’s roadmap around a forthcoming Llama is describing a product line the company is no longer shipping under that name.
| Primitive | Weights published | Permissive licence | Commercial use unrestricted |
|---|---|---|---|
| Muse Glimmer | Hugging Face | Apache 2.0 | No threshold |
| GLM-5.2 | Hugging Face | MIT | No threshold |
| LongCat-Flash-Lite-Sparse | Hugging Face | MIT | No threshold |
| Qwen3.8-Max | Hugging Face | Custom licence | Gated above $50M |
| Kimi K3 | Hugging Face | Custom licence | Gated above $20M |
Grok 4.6: same sticker, more expensive cache (August 12)
xAI released Grok 4.6 on August 12, 35 days after Grok 4.5. The headline rate is unchanged at $2 input and $6 output per million tokens with a 500K context window, and there is a fast variant at double the price. xAI’s stated benchmarks are an Artificial Analysis Intelligence Index of 61, GDPVal-AA v2 of 1753, CursorBench v3.2 at 69.9%, DeepSWE v1.1 at 65.9% and FrontierCode v1.1 Extended at 61.3%. Artificial Analysis independently confirms the 61 at high effort, which ties GPT-5.6 Sol and trails only Claude Opus 5 at 63 and Claude Fable 5 at 62.
The detail the launch post does not lead with sits in xAI’s model documentation: cached input rose from $0.30 per million tokens on Grok 4.5 to $0.50 on Grok 4.6. For an agent loop that replays a large cached prefix on every iteration, cached input is often the dominant line on the bill, so a 67% rise there can outweigh an unchanged sticker rate. Both models also double every rate once a prompt reaches 200K tokens, and the higher rate then applies to every token in the request, not just the ones past the threshold.
Gemini 3.7 Flash: five points of index for half the price (August 13)
Google announced Gemini 3.7 Flash on August 13 across Google AI Studio, Android Studio, the Gemini Enterprise Agent Platform and Gemini Spark for AI Pro and Ultra subscribers in more than 160 countries. Google’s own comparison against Gemini 3.6 Flash: DeepSWE v1.1 up to 65.3% from 49.0%, FrontierCode 1.1 Main to 43.6% from 34.4%, WebDev Arena Elo to 1588 from 1538, GDP.pdf to 34.0% from 22.0% and AutomationBench to 30.4% from 17.0%.
Independently, Artificial Analysis scores it 56 on the Intelligence Index at high effort. That is the number worth holding onto, because at $0.75/$3.75 it puts a model four points below Grok 4.6 and five below GPT-5.6 Sol at a fraction of their token price, at least until the end of December. Bloomberg reported the launch as arriving while Google’s top-end Gemini 3.5 Pro remains delayed, which has now been the case since Gemini 3.5 Flash shipped in May.
GLM-5.3: the same base model, post-trained harder (August 14)
Z.ai released GLM-5.3 on August 14, and it is the most interesting engineering story of the month even though it is the thinnest on published economics. It is not a new base model. Z.ai states it runs on the same base as GLM-5.2 and that every capability gain comes from scaled post-training: more task environments, more environment types and longer training. Z.ai’s model documentation confirms a 1 million-token context window and 128K max output.
The vendor-reported jumps are concentrated exactly where a post-training story would predict, on long-horizon agentic work rather than single-turn quality. Terminal-Bench 3.0 goes from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents’ Last Exam on the CLI from 23.8 to 28.5. On the security side Z.ai reports CyberGym at 84.5% from 77.2% and ExploitBench at 54.4% from 24.4%, and says the cyber capability scaled faster than it expected. All of those are Z.ai’s own numbers on Z.ai’s own runs, and none has been independently reproduced.
Two caveats mattered more than the benchmark table at launch. One has since resolved. The first was that no price was published: for four days the model existed and its cost did not. Z.ai’s pricing page now lists GLM-5.3 at $1.40 input, $0.26 cached input and $4.40 output per million tokens, checked August 21. That is not a new number. It is the number GLM-5.2 already carried, and GLM-5.1 before it.
The second caveat still stands. Unusually for Z.ai, the weights were not published at launch. GLM-5.2 shipped MIT-licensed and downloadable; GLM-5.3 remains API-only, with Z.ai saying it will publish weights roughly two weeks after launch once safety evaluation and hardening finish, and naming no licence in advance.
That last point is the one to hold. A model that is announced as open, benchmarked as open, and covered as open, but is not yet downloadable, was the same pattern as Qwen3.8-27B and Muse Spark 1.2 earlier in the month. Alibaba has since closed its half of that gap, publishing the Qwen3.8-27B weights on August 14 under Apache 2.0. Meta and Z.ai have not. If the GLM-5.3 weights land on schedule with an MIT licence, it becomes the strongest open coding model on these numbers. Until then it is a closed model with an open-weights promise attached.
The second half: a rate card, a cloaked listing, and the two models that ended the month (August 15 to 31)
No frontier model shipped between August 15 and August 25. Four things happened anyway, and the last of them reframed the first.
Z.ai priced GLM-5.3, and priced it at zero premium. The rate card that was missing at launch now reads $1.40 input, $0.26 cached input and $4.40 output per million tokens: identical to GLM-5.2, and identical to GLM-5.1. Three consecutive generations at one price. Read that against the rest of the month and it is the outlier position. Google cut its Flash tier on a timer that expires. DeepSeek split one flat rate into two rates that depend on the hour. Z.ai shipped a model it claims is six times better at Terminal-Bench 3.0 and left the price alone.
For a buyer, the flat rate is worth more than it looks. A price that does not move between generations is the only one of the three that can be put in a twelve-month budget without a footnote. It also means the upgrade decision from GLM-5.2 to GLM-5.3 carries no cost argument at all, in either direction: the same tokens cost the same money, so the only question is whether the post-training gains are real, and that still rests on Z.ai’s own unreproduced runs.
Tencent shipped two translation models on August 20. Hy-MT2-30B-A3B is the flagship, a mixture-of-experts model with about 3 billion active parameters, and Hy-MT2-1.8B is the compact sibling. Both cover 33 language pairs plus five Chinese dialect and minority-language pairs, and both are listed at an 8,000-token context window: $0.074 input and $0.295 output per million tokens for the 30B, and $0.044 and $0.177 for the 1.8B, as listed on OpenRouter. The context window is the detail worth reading. In a month where three separate releases advertised a million tokens, shipping 8K is a deliberate statement that this is a sentence-level tool and not an agent, priced accordingly at roughly a twentieth of what the cheapest general model on this page charges for output.
A cloaked model appeared on August 20, and on August 26 it turned out to be Z.ai again. A listing called Ox Alpha went up on OpenRouter with no named maker, attributed only to “Stealth”, free while it was being evaluated, with a 1.05 million-token context window. Six days later Z.ai claimed it: Ox Alpha was GLM-5.3-Flash. The Ox Alpha explainer covers how the listing was read before the reveal.
The reveal is the month’s best price story, and it undercuts the flat-rate reading above. GLM-5.3-Flash is a 320B-total, 18B-active mixture-of-experts model on a hybrid sparse-plus-linear attention design, the first natively multimodal model in the GLM-5 line, with a 1,048,576-token context window. It shipped MIT-licensed with weights on Hugging Face, and Z.ai lists it at $0.15 input and $0.50 output per million tokens, with a launch promotion halving that to $0.075/$0.25 through September 9. The independent Artificial Analysis Intelligence Index scores it 57, above the 56 it gives Gemini 3.7 Flash.
So Z.ai did not hold its price flat across the month. It held GLM-5.3 flat at $1.40/$4.40 and then, twelve days later, shipped a higher-scoring model at roughly a ninth of that rate. Anyone who read the GLM-5.3 rate card on August 21 as Z.ai’s position on pricing read it wrong by twelve days.
Alibaba closed the month with an architecture preview. Qwen3.8-Flash-Next landed the same day, August 26, and it is not a normal point release: Alibaba describes it as an experimental preview of the architecture that will underpin Qwen4. The model card puts it at 125B total parameters with 6B activated, plus a 51B n-gram embedding layer and 4B multi-token prediction, on a hybrid of Gated DeltaNet and a new Qwen Sparse Attention that operates on micro-blocks rather than individual tokens. Context is 262,144 native, extensible to a million. Weights are open under the Qwen Community 1.0 licence, and Qwen Cloud prices the hosted version at $0.15 input and $0.47 output per million tokens.
Two labs, one day, and the same move: put a frontier-adjacent model at roughly a sixth to a ninth of the flagship rate and give the weights away. That, not the mid-month rate cards, is what August actually settled.
DeepSeek shipped an experimental vision model on August 21. DeepSeek-V4-Flash-Vision-Exp is the lab’s first multimodal vision understanding model, reached by setting model='deepseek-v4-flash-vision-exp', and it carries exactly the V4 Flash rate card: $0.44 input and $1.32 output per million tokens at peak, halved off-peak, with images converted to input tokens by their dimensions. Charging nothing extra for vision is the notable part; the experimental label is the reason it is not counted as a flagship here.
Two smaller changes also landed in the window: Grok 4.6 became available on Amazon Bedrock on August 19, a week after its launch, which moves who can buy it rather than what it costs. And on August 27 Z.ai finally published the GLM-5.3 weights it had promised for the week of August 24, roughly on schedule but not under MIT: the repository carries an “other” licence, unlike GLM-5.2, which shipped MIT. The commitment was met; the licence quietly was not the one its predecessor set.
Did Anthropic release a new model in August 2026?
No. Anthropic’s newest model through August was Claude Opus 5, released July 24 at $5/$25 per million tokens, which led the Artificial Analysis Intelligence Index at 63 on max effort. The drought ended the day after this window closed: Claude Fable 5.1 shipped on September 1, holding Fable 5 rates and cutting only the cache read. Within August itself, two things changed on the commercial side that matter more than a release would have.
The first is the Sonnet 5 increase being cancelled, covered above. The second is that Claude Opus 4.1 was retired on August 5, having been deprecated on June 5, though Anthropic’s pricing page still lists it as available on Amazon Bedrock and Google Cloud. Its $15/$75 rate is a useful marker of how far flagship pricing has moved: Claude Opus 5, which supersedes it and outscores it comfortably, costs a third as much. One migration trap worth flagging: temperature, top_p and top_k now return a 400 error on Claude Opus 4.7 and later, so code written against 4.1 needs editing rather than simply repointing at a new model id. Our Opus 5 pricing and benchmarks breakdown has the full picture, and the Opus 5 versus Sonnet 5 comparison covers which tier to actually run.
Did OpenAI release a new model in August 2026?
No new model. OpenAI’s most recent release remains the GPT-5.6 family (Sol, Terra and Luna), generally available since July 9, whose prices it cut on July 30, taking Luna down 80% to $0.20/$1.20 and Terra down 20% to $2/$12. On August 21 flagship Sol, which had been excluded from that cut, took its own promotional reduction to $4/$20 per million tokens, in effect through at least November 21, 2026. Sol scores 61 on the Artificial Analysis Intelligence Index at max effort, tying Grok 4.6.
The change in August was distribution rather than capability: GPT-5.6 Sol became the default model for Plus and Pro users, and Luna the default for Free and Go, as reported in early August. That matters commercially because it moves the majority of ChatGPT traffic onto models OpenAI has already made much cheaper to serve.
Which August model should you actually use?
Based on the verified rates and the independent index scores above, not the launch tables:
- Cheapest capable tier, right now: GLM-5.3-Flash at $0.15/$0.50 with an index of 57, which arrived on August 26 and displaced the answer this section carried for most of the month. It scores a point above Gemini 3.7 Flash and costs a fifth as much, and the weights are MIT, so the hosted rate is a ceiling rather than a lock-in. Gemini 3.7 Flash at $0.75/$3.75 remains the pick if you need Google’s stack specifically, but put a January 1 reminder in your budget, because that rate doubles to $1.50/$7.50.
- Cheapest capable tier you can self-host: Qwen3.8-Flash-Next, 125B total with only 6B active, at $0.15/$0.47 hosted. Read the Qwen Community 1.0 licence before building on it, and treat it as what Alibaba calls it: an experimental preview of the Qwen4 architecture, not a settled production model.
- Top of the board: Claude Opus 5 at 63, then Claude Fable 5 at 62, then Grok 4.6 and GPT-5.6 Sol tied at 61. Grok 4.6 is by far the cheapest of those four at $2/$6, so it is the one to price first if your work is agentic and cache-heavy, with the caveat that its cached-input rate rose.
- Self-hosting with commercial intent: Muse Glimmer if the licence matters most, since Apache 2.0 with no revenue gate is the cleanest position of any model here, and it runs on one consumer GPU. Qwen3.8-Max if capability matters most, but read the $50M clause before you build a product on it.
- Long-horizon agent work: GLM-5.3 is priceable at $1.40/$4.40, the same rate as GLM-5.2, and its weights landed on August 27 under an “other” licence. The cost objection is gone, but so is the reason to reach for it first: GLM-5.3-Flash scores 57 on the independent index at a ninth of the price. Benchmark both on your own harness before assuming the expensive one wins, because the Terminal-Bench 3.0 jump from 4.6 to 28.3 that made GLM-5.3 interesting is still Z.ai grading its own work.
- Anyone currently on DeepSeek: the peak and off-peak split took effect on August 16, so this has already hit live bills. The DeepSeek Harness breakdown has the per-line-item multiples, including the cached-input rate that rose hardest.
- Anyone who planned a September migration off Claude Sonnet 5: cancel it. The increase you were avoiding is not happening.
For a head-to-head on your own workload, the model comparison tool prices two or three models against the same task, the cost-per-task calculator takes your own token profile, and the value leaderboard ranks the field by benchmark points per dollar.
The release timeline from here
What is confirmed, announced or credibly rumored as of August 31, 2026:
- Gemini 3.5 Pro (Google DeepMind): announced, still no date. It has been listed as coming soon since May, and Bloomberg framed the August 13 Flash launch as arriving while the top-end model remains delayed. Google is now the only top-three lab that has not shipped a frontier-tier model since May.
- Qwen3.8-27B (Alibaba): delivered. Alibaba pointed to the week of August 10 for the smaller sibling of Qwen3.8-Max, and the weights landed on August 14. The model card shows a dense 27B native vision-language model under Apache 2.0, with a 262,144-token native context extensible to about a million, and no revenue gate of the kind attached to Qwen3.8-Max. This is the month’s one clean case of an open-weights commitment being met on time and under a real open-source licence.
- Muse Spark 1.2 open weights (Meta): announced, unpublished. The API model shipped August 5; the weights have not appeared.
- GLM-5.3 weights (Z.ai): delivered on August 27, but not under MIT. Z.ai committed to roughly two weeks after the August 14 launch and made it, publishing the weights to zai-org/GLM-5.3 on August 27. The licence is the catch: the repository is tagged “other”, not the MIT that GLM-5.2 shipped under. GLM-5.3-Flash, released two days earlier, did get MIT. Read the licence per model, not per lab.
- GPT-6 (OpenAI): rumor only. Still no model page, no date and no confirmation of the name. (Resolved after this month closed: OpenAI released GPT-6 Astra on September 3, 2026.)
- Ox Alpha: resolved. The cloaked OpenRouter listing that went live on August 20 was Z.ai, and the model was GLM-5.3-Flash, announced on August 26. Six days from stealth listing to named release, which is at the fast end of the usual range.
The pattern worth reading is that four separate commitments to publish open weights are still outstanding, across Meta, Z.ai, Ant Group and Black Forest Labs. Announced is not shipped, and a release tracker that counts announcements rather than artifacts will overstate the month. Two labs have now delivered: Meta with Muse Glimmer on the day, under Apache 2.0, and Alibaba with Qwen3.8-27B on August 14, also under Apache 2.0. Both of the licences that actually arrived were permissive; both of the licences still being promised are unnamed.
Frequently asked questions
- What new AI models came out in August 2026?
- Twelve distinct models: Qwen3.8-Max from Alibaba on August 3, Muse Spark 1.2 from Meta on August 5, Muse Glimmer from Meta on August 10, Grok 4.6 from xAI on August 12, Gemini 3.7 Flash from Google on August 13, GLM-5.3 from Z.ai and Qwen3.8-27B from Alibaba on August 14, the Hy-MT2-30B-A3B and Hy-MT2-1.8B translation models from Tencent on August 20, DeepSeek V4 Flash Vision Exp on August 21, and GLM-5.3-Flash from Z.ai and Qwen3.8-Flash-Next from Alibaba on August 26. Five of the twelve shipped with downloadable weights.
- What is the latest AI model right now?
- As of August 31, 2026 the newest releases are GLM-5.3-Flash from Z.ai and Qwen3.8-Flash-Next from Alibaba, both on August 26. GLM-5.3-Flash is the model that had been listed anonymously as Ox Alpha since August 20: a 320B-total, 18B-active multimodal mixture-of-experts model, MIT-licensed, scoring 57 on the Artificial Analysis Intelligence Index at $0.15 input and $0.50 output per million tokens.
- Is GLM-5.3 open source?
- The weights are downloadable, but the licence is not MIT. GLM-5.3 was API-only at its August 14 launch; Z.ai published the weights on August 27, roughly the two weeks it had promised, under an "other" licence rather than the MIT that GLM-5.2 shipped under. Its cheaper sibling GLM-5.3-Flash, released August 26, did ship MIT-licensed.
- Did Anthropic release a new model in August 2026?
- No. Its newest model through August was Claude Opus 5, released July 24; Claude Fable 5.1 followed on September 1, just outside this window. Anthropic did make two commercial changes in August: it cancelled the scheduled September 1 increase on Claude Sonnet 5, making the $2/$10 per million token rate permanent, and it retired Claude Opus 4.1 on the Claude API on August 5.
- Did OpenAI release a new model in August 2026?
- No. The most recent OpenAI models remain GPT-5.6 Sol, Terra and Luna, generally available since July 9 and repriced on July 30. In August OpenAI made Sol the default for Plus and Pro users and Luna the default for Free and Go.
- Is DeepSeek raising its prices?
- Yes. From 16:00 UTC on August 16, 2026 DeepSeek moves to peak and off-peak billing. For DeepSeek V4 Pro the output rate becomes $1.98 per million tokens off-peak and $3.96 at peak, against $0.87 today, so even the cheapest hour costs 2.3 times the current price. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC.
- Is Qwen3.8-Max open source?
- The weights are downloadable, but the licence is not an open-source licence. Qwen3.8-Max ships under a custom qwen3.8-max licence rather than Apache 2.0. It requires prominent attribution above 100 million monthly active users or US$20 million monthly revenue, and requires a separate licence from Qwen once a model-as-a-service or AI work-assistant business exceeds US$50 million of revenue over twelve consecutive months. Internal use is exempt.
- How much does GLM-5.3 cost?
- Z.ai lists GLM-5.3 at $1.40 per million input tokens, $0.26 per million cached input tokens and $4.40 per million output tokens. No price was published at the August 14 launch; the rate appeared on the pricing page within the following week. It is the same rate Z.ai charges for GLM-5.2 and GLM-5.1, so the newer model carries no price premium over either predecessor.
- What is the cheapest new AI model in August 2026?
- Among general-purpose models, Gemini 3.7 Flash at $0.75 input and $3.75 output per million tokens, which scores 56 on the independent Artificial Analysis Intelligence Index. That rate is introductory: Google states it doubles to $1.50 and $7.50 on January 1, 2027. Cheaper still, but only for one job, are the Tencent Hy-MT2 translation models at $0.044 and $0.074 input.
- What AI models are coming next after August 2026?
- Gemini 3.5 Pro from Google DeepMind remains announced with no date and is now several months delayed. Z.ai committed to GLM-5.3 weights roughly two weeks after its August 14 launch, pointing at the week of August 24. Meta announced open weights for Muse Spark 1.2 that have not been published. A cloaked coding model called Ox Alpha has been serving free traffic on OpenRouter since August 20 with no lab named. GPT-6 remained a rumor with no model page or date through August; it was released as GPT-6 Astra on September 3, 2026. Alibaba has already delivered on Qwen3.8-27B, which shipped on August 14 under Apache 2.0.
Sources
- Alibaba Qwen (2026). Qwen3.8-2.4T-A95B model card and licence. Hugging Face. Parameters, context window, benchmarks and the custom licence terms. Verified August 14, 2026.
- Anthropic (2026). Pricing. Claude platform documentation. Claude Sonnet 5 standard rate and the cancelled September 1 increase; Opus 5 and Opus 4.1 rates. Verified August 14, 2026.
- Anthropic (2026). Model deprecations. Claude platform documentation. Claude Opus 4.1 retirement date and parameter deprecations. Verified August 14, 2026.
- DeepSeek (2026). Models and pricing. DeepSeek API documentation. Current and post-August 16 peak/off-peak rates for V4 Pro and V4 Flash. Verified August 14, 2026.
- DeepSeek (2026). News and updates. DeepSeek API documentation. V4 Pro general availability and the pricing change effective date. Verified August 14, 2026.
- Google (2026). Introducing Gemini 3.7 Flash. Google blog. Launch date, availability surfaces and vendor benchmark comparison. Verified August 14, 2026.
- Google (2026). Gemini API pricing. Google AI for Developers. Gemini 3.7 and 3.6 Flash rates and the January 1, 2027 increase. Verified August 14, 2026.
- Meta (2026). Muse-Glimmer-30B model card. Hugging Face. Apache 2.0 licence, parameters, context, knowledge cutoff and vendor benchmarks. Verified August 14, 2026.
- Meta (2026). Muse Spark. Meta developer documentation. Muse Spark 1.2 per-token rates, contributor tier and context window. Verified August 14, 2026.
- xAI (2026). Grok 4.6. xAI news. Release date, headline pricing and vendor benchmark scores. Verified August 14, 2026.
- xAI (2026). Models. xAI documentation. Cached-input rates and long-context pricing for Grok 4.5 and 4.6. Verified August 14, 2026.
- Artificial Analysis (2026). Model leaderboards. Independent model evaluations. Intelligence Index scores with reasoning-effort variants. Read August 14, 2026.
- Meituan (2026). LongCat-Flash-Lite-Sparse model card. Hugging Face. MIT licence and parameter counts. Verified August 14, 2026.
- Moonshot AI (2026). Kimi K3 licence. Hugging Face. The US$20 million revenue threshold that triggers a separate commercial agreement. Verified August 14, 2026.
- Z.ai (2026). GLM-5.2 model card. Hugging Face. MIT licence, 753B parameters and 1M context. Verified August 14, 2026.
- Z.ai (2026). GLM-5.3. Z.ai model documentation. Context window and max output. Verified August 14, 2026.
- Z.ai (2026). Pricing. Z.ai documentation. GLM-5.3 at $1.40 input, $0.26 cached and $4.40 output per million tokens, identical to GLM-5.2 and GLM-5.1. No GLM-5.3 rate was listed at launch. Verified August 21, 2026.
- Alibaba Qwen (2026). Qwen3.8-27B model card. Hugging Face. Apache 2.0 licence, dense 27B native vision-language model, 262,144-token native context extensible to about one million. Verified August 21, 2026.
- Z.ai (2026). GLM-5.3-Flash. Z.ai documentation. GLM-5.3-Flash, released August 26 as the reveal of the Ox Alpha stealth listing, 1M-token context. Standard rate $0.15 input and $0.50 output per million tokens on the Z.ai pricing page, halved to $0.075/$0.25 by a launch promotion running to September 9, 2026. Verified August 31, 2026.
- Z.ai (2026). zai-org/GLM-5.3. Hugging Face. GLM-5.3 weights published August 27, 2026 under an “other” licence rather than the MIT that GLM-5.2 carries. Supports the licence claim in the release timeline. Verified August 31, 2026.
- Alibaba Qwen (2026). Qwen3.8-Flash-Next model card. Hugging Face. Released August 26, 2026: 125B total parameters with 6B activated plus a 51B n-gram embedding layer, Gated DeltaNet and Qwen Sparse Attention, 262,144-token native context extensible to one million, Qwen Community 1.0 licence. Verified August 31, 2026.
- Qwen Cloud (2026). Qwen3.8-Flash overview. Alibaba Qwen Cloud. Hosted pricing of $0.15 input and $0.47 output per million tokens, and the million-token context window. Verified August 31, 2026.
- DeepSeek (2026). News and updates. DeepSeek API documentation. DeepSeek-V4-Flash-Vision-Exp released August 21, 2026 as the lab first multimodal vision understanding model, priced identically to V4 Flash at $0.44/$1.32 per million tokens at peak and half that off-peak, with images billed as input tokens by dimension. Verified August 31, 2026.
- OpenRouter (2026). Models, newest first. OpenRouter model directory (aggregator, not the maker). Listing dates and rates for Tencent Hy-MT2-30B-A3B and Hy-MT2-1.8B on August 20, and the unattributed Ox Alpha listing. Used because Tencent publishes no English rate card for these models. Verified August 21, 2026.
- Asif Razzaq (2026, August 14). Z.ai ships GLM-5.3 without retraining the base model. MarkTechPost (secondary). Release date and the vendor-reported benchmark deltas against GLM-5.2, which Z.ai’s own client-rendered blog post could not be read directly to confirm. Verified August 14, 2026.
- Capital & Compute. New AI models released in July 2026. The previous installment of this series.