New AI Models Released in October 2026: Prices
Every AI model released in October 2026, with dates, verified launch prices and a primary source for each. Updated through the month as releases land.

Ten AI models were released in the first four days of October 2026, from seven labs, and all ten landed on October 1. None came from OpenAI, Anthropic, Google, xAI, Meta or DeepSeek. The headline is a new kind of model: Cloudflare, AWS Strands Labs and Perplexity each shipped an open decision model that answers with probabilities instead of text, bills input only and charges nothing for output. The cheapest is Perplexity’s pplx-decider-v1-27b at $0.04 per million input tokens. The same day brought Microsoft’s MAI speech trio, Black Forest Labs’ FLUX 3 Image, a gated Tavus video preview, and the month’s only price cut so far, Unbiased’s Pareto 26.10 Preview.
The quick verdict for anyone choosing today: for a high-volume yes or no, label or routing call, pplx-decider-v1-27b is the cheapest managed option and is open-weight. Clef-flash is the one to pick when latency matters more than price. For general-purpose text work, nothing launched this month replaces a September model yet.
Every AI model released in October 2026
| Model | Provider | Released | Launch price | What it is |
|---|---|---|---|---|
| pplx-decider-v1-27b | Perplexity | Oct 1 | $0.04 input / $0 output per Mtok | Open decision model behind the Perplexity Decisions API. Apache 2.0 |
| Clef | Cloudflare | Oct 1 | $0.24 input / $0 output per Mtok | Cloudflare’s first in-house model. Open decision model on Workers AI |
| Clef-flash | Cloudflare | Oct 1 | $0.09 input / $0 output per Mtok | Latency tier of Clef: 38.8 ms median decision. Apache 2.0 |
| Strands Decider 2B | AWS Strands Labs | Oct 1 | Self-hosted, no rate | 2B open decision model that runs locally, training recipe published |
| Pareto 26.10 Preview | Unbiased | Oct 1 | $0.80 / $3.20 per Mtok | Preview of the next Pareto blend, 68% below Pareto 26.9 on input |
| MAI-Transcribe-2-Streaming | Microsoft | Oct 1 | $0.54 per audio hour (introductory) | Microsoft’s first streaming speech-to-text model |
| MAI-Voice-2.1 | Microsoft | Oct 1 | $22 per million characters | Text-to-speech in 23 languages and 26 locales |
| MAI-Voice-2.1-Flash | Microsoft | Oct 1 | $15 per million characters | Low-latency text-to-speech tier, about 150 ms end to end |
| FLUX 3 Image | Black Forest Labs | Oct 1 | $0.041 to $0.607 per image | Image generation and editing half of FLUX 3, up to 4K and ten references |
| Griffin-Lite | Tavus | Oct 1 | Not published | Full-duplex video-to-video research preview, invited testers only |
What is a decision model, and why did three ship on one day?
A decision model does not write. You send it the state of your application and a typed question (yes or no, pick one of these options, place this on a scale), and it returns one answer with a probability attached to every option. There is no free text to parse, so the output cannot drift out of format, and the probability lets you set a threshold and send the unsure cases to a person or a larger model.
The category was defined by TypeSafe’s Jev, which charges $0.042 per million input tokens and nothing for output; the full breakdown is in what Jev is and what it costs. October 1 is the day the format stopped being one company’s product. Three open alternatives shipped within hours of each other:
- Clef and Clef-flash, Cloudflare’s first models trained in-house, announced in the Cloudflare blog post Clef: open decision models under Apache 2.0. Cloudflare reports a median decision time of 209.3 ms for Clef and 38.8 ms for Clef-flash, with 238.6 ms and 122.4 ms at the 95th percentile. Both take a 65,536-token context, per the Workers AI model documentation.
- Strands Decider 2B, from AWS’s Strands Labs, described in Introducing Strands Decider 2B. It takes a Qwen 2B base and replaces the next-word head with a pointer head of just over a million parameters that scores your options. Strands reports a median of about 115 ms on an Nvidia RTX 3090 and 153 ms on a MacBook M3, and ships the training data and scripts alongside the weights.
- pplx-decider-v1-27b, Perplexity’s model behind its new Decisions API, fine-tuned from Qwen3.8-27B and published on Hugging Face under Apache 2.0. Perplexity’s own model card reports 85.71% across 11 benchmarks, against 74.76% for the base model it started from. That is a vendor-run comparison, not an independent one.
The wave did not start on October 1. Three closed decision models launched in the last two days of September: Liquid AI’s d1 on September 29 (as reported by MarkTechPost), and Inception’s Mercury Decide and Fastino’s GLiDE on September 30. They are September releases by their own dates, which is why they are not in the October table. What October 1 added was open weights: all three October decision models can be downloaded and run on your own hardware.
Is a decision model cheaper than a small chat model?
Only some of them, and the free output matters less than it sounds. A chat model asked for a one-word label writes very little, so its output bill is small. The input rate decides the cost, and on input, Clef is more expensive than the cheapest chat tier on the market.
| Scenario | Input | Output | Total |
|---|---|---|---|
| pplx-decider-v1-27b | $0.080 | $0.000 | $0.080 |
| Jev 1.13.0 | $0.084 | $0.000 | $0.084 |
| Clef-flash | $0.180 | $0.000 | $0.180 |
| GPT-6 Luna, short label | $0.200 | $0.015 | $0.215 |
| GPT-6 Luna, with reasoning | $0.200 | $0.265 | $0.465 |
| Clef | $0.480 | $0.000 | $0.480 |
The model behind the chart is simple and every input is a list price. Each decision sends 2,000 tokens of state and question. GPT-6 Luna, OpenAI’s cheapest tier at $0.10 input and $0.50 output per million tokens on the OpenAI API pricing page, is costed two ways: answering with a 30-token label, and answering after about 500 tokens of reasoning.
Three things follow from it:
- pplx-decider and Jev are the only clear price wins. At $0.04 and $0.042 per million input tokens they cost less than half of a short-label Luna call, and the gap widens as soon as the chat model reasons.
- Clef-flash roughly ties a short-label Luna call ($0.180 against $0.215). Its case is speed and the probability on every answer, not a lower bill.
- Clef costs more than a short-label Luna call and about the same as a Luna call that reasons. Pick it for the calibrated probabilities, the open weights and staying inside Cloudflare’s network, not for price.
The chart leaves out what a decision model does not do: it cannot explain itself, write a reply or call a tool. The comparison holds only for jobs that end in a choice.
Which October decision model should you pick?
- Cheapest managed calls: pplx-decider-v1-27b at $0.04 per million input tokens, with images accepted alongside text.
- Lowest latency on a hosted API: Clef-flash, at a 38.8 ms median by Cloudflare’s measurement, already inside Workers AI if your app runs there.
- On your own hardware or offline: Strands Decider 2B. At 2B parameters it runs on a laptop GPU, and Strands ships the training data and scripts with the weights.
- Already on Jev: there is no price reason to move. Jev’s $0.042 is within a fraction of a cent of pplx-decider per thousand calls, so test the accuracy on your own questions before switching.
How to get started with a decision model
- Pick five to ten decisions your app already makes with a chat model and a parsing step: ticket routing, moderation, a yes or no gate before a tool call.
- Log the inputs and the chat model’s answers for a few hundred real cases. That log is your test set.
- Run the same cases through pplx-decider-v1-27b or Clef-flash and compare agreement and the probability spread. Set a threshold below which the case falls back to the chat model.
- Price the switch with the chart’s method: your real input tokens per call times the input rate, against your current per-call bill.
Pareto 26.10 Preview: the first price cut of October
Two weeks after Pareto 26.9 launched at $2.50 and $7.50 per million tokens, Unbiased published a preview of its successor at $0.80 input, $0.03 cached input and $3.20 output, according to the Unbiased model card. That is 68% less on input and 57% less on output than the model it previews. The only benchmark published so far is Unbiased’s own, preliminary and vendor-run: 69.9% on DeepSWE v1.1 at $0.24 per task. Unbiased says the figures may change before final publication and that its cost method still needs confirming.
Treat the price as a preview price. Unbiased’s own listing on OpenRouter describes the model as a preview of the next Pareto version and points stable workloads at Pareto 26.9.
Microsoft MAI speech models: what does $0.54 an hour buy?
Microsoft shipped three speech models on October 1, all generally available, per its announcement Our first streaming transcription model debuts at no. 1 on Artificial Analysis:
- MAI-Transcribe-2-Streaming transcribes speech as it arrives, in 60 languages. Microsoft cites the independent Artificial Analysis AA-WER Streaming evaluation, where it ranks first at a 2.5% word error rate with the final transcript ready 0.13 seconds after the speaker stops. It costs $0.54 per audio hour at an introductory rate that runs through the end of 2026.
- MAI-Voice-2.1 is text-to-speech across 23 languages and 26 locales at $22 per million characters, and keeps a single voice natively accented in every language.
- MAI-Voice-2.1-Flash generates 45 seconds of audio at about 150 ms end-to-end latency for $15 per million characters, 32% less than the full model.
The streaming transcriber is the one to watch for cost. A voice agent pays for every second of audio the caller speaks, so the year-end expiry of the introductory rate is a date to put in the budget now, not in January.
FLUX 3 Image: what does an image cost?
Black Forest Labs announced FLUX 3 in July as one network for image, video, audio and robot actions. FLUX 3 Image is the image half reaching the API, dated October 1 in the Black Forest Labs release notes. It takes up to ten reference images, places elements by bounding box and renders in 15 aspect ratios. Pricing is per image, by resolution:
| Resolution | Price per image |
|---|---|
| 768 square | $0.041 |
| 1K | $0.048 |
| 2K | $0.100 |
| 4K | $0.607 |
The step to 4K is the one to notice: a 4K image costs six times a 2K one.
Tavus Griffin-Lite: a video model you cannot buy yet
Tavus describes Griffin as a full-duplex video-to-video model: it watches, listens and replies in speech and video at once, rather than passing a call through separate speech, language and video systems. Only the smaller Griffin-Lite is out, as a research preview for invited testers, per the Tavus Griffin page. The full model is held back until safety concerns are addressed. Tavus reports a score of 3.83 out of 5 on NVIDIA’s VideoFDB generation benchmark, against a human reference of 3.92 and 2.80 for the next system. In a company-run test, 26 of 54 people (48%) believed after a one-minute call that they had been talking to a real person. No price is published.
The release pace
October has ten releases in four days, and all ten share one date. For comparison, the release archive counts 5 tracked releases in June (from June 9, when the record starts), 21 in July, 12 in August and 34 in September. A ten-release first day says little about the month’s total. It does say that October opened without a frontier launch.
Several models appear on other October lists but not this one, because their own announcements date them elsewhere. Liquid d1, Mercury Voice and Mercury Decide, GLiDE, Ant Group’s Ling-3.1-flash and Bilibili’s Index-Translate all launched on September 29 or 30. Apodex 1.1 Mini appears as an October 1 release on one gateway, but its weights shipped with Apodex 1.1 at the start of September; October 1 is when the gateway listed it. GPT-Synopsys, announced by OpenAI and Synopsys on September 30, is a partnership to build a chip-design model, not a model anyone can use yet. Every date on this page is read off the lab’s own announcement or documentation. The method is set out in the AI model release archive.
Did OpenAI release a new model in October 2026?
Not as of October 4, 2026. OpenAI’s most recent release is GPT-6.1 Sol, shipped at DevDay on September 29 at $2 input and $10 output per million tokens with the cache read halved to $0.10. Its cheapest current tier, GPT-6 Luna at $0.10 and $0.50, launched on September 22. Both are covered in the September roundup, and current rates for every OpenAI model are in the AI model tracker.
Did Google or Anthropic release a new model in October 2026?
Not yet. Google’s latest is Gemini 4 Argon, announced on September 30 at an introductory $2 and $10 per million tokens and still limited to cyber defenders in its Fairwind Program; the access and pricing detail is in the Gemini 4 Argon breakdown. Anthropic’s latest is Claude Sonnet 5.5, released September 28 on Sonnet 5’s unchanged $2 and $10 rate card, six days after Claude Opus 5.5.
Frequently asked questions
- What AI models were released in October 2026?
- As of October 4, 2026, ten, all on October 1. Perplexity released pplx-decider-v1-27b at $0.04 per million input tokens with free output; Cloudflare released Clef at $0.24 and Clef-flash at $0.09, also with free output; AWS Strands Labs released the self-hosted Strands Decider 2B. Unbiased released Pareto 26.10 Preview at $0.80 input and $3.20 output. Microsoft released MAI-Transcribe-2-Streaming at $0.54 per audio hour, MAI-Voice-2.1 at $22 per million characters and MAI-Voice-2.1-Flash at $15. Black Forest Labs released FLUX 3 Image at $0.041 to $0.607 per image, and Tavus opened Griffin-Lite as an invite-only research preview with no price. No OpenAI, Anthropic, Google, xAI, Meta or DeepSeek model has launched in October so far.
- What is the cheapest decision model in October 2026?
- Perplexity pplx-decider-v1-27b, at $0.04 per million input tokens with output free, just under the $0.042 that TypeSafe Jev charges. Cloudflare Clef-flash costs $0.09 and Clef $0.24 per million input tokens, also with free output. Strands Decider 2B has no hosted rate; you run it on your own hardware.
- Is a decision model cheaper than GPT-6 Luna?
- Only pplx-decider and Jev are clearly cheaper. On a modeled 1,000 decisions at 2,000 input tokens each, pplx-decider costs $0.080 and Jev $0.084, against $0.215 for GPT-6 Luna answering with a short label. Clef-flash costs $0.180, roughly a tie, and Clef costs $0.480, more than a short-label Luna call and about the same as a Luna call with 500 reasoning tokens.
- Did any AI model get cheaper in October 2026?
- One so far. Unbiased Pareto 26.10 Preview, released October 1, lists at $0.80 input and $3.20 output per million tokens, 68% and 57% below the $2.50 and $7.50 that Pareto 26.9 launched at on September 17. It is a preview, so the rate and the scores may still change.
- How many AI models has the release archive tracked through October 2026?
- Eighty-two dated launches from June 9, 2026: 5 in June, 21 in July, 12 in August, 34 in September and 10 in October as of October 4. Every row carries the price the model launched at and a link to the provider primary source.
Sources
Cloudflare (2026). Clef: open decision models. Cloudflare blog, October 1, 2026: Apache 2.0, 64K context, vendor-measured median and p95 decision times. https://blog.cloudflare.com/clef-decision-models/ Verified 2026-10-04.
Cloudflare (2026). Clef and Clef-flash. Workers AI model documentation: $0.24 and $0.09 per million input tokens, 65,536-token context. https://developers.cloudflare.com/workers-ai/models/clef/ and https://developers.cloudflare.com/workers-ai/models/clef-flash/ Verified 2026-10-04.
AWS Strands Labs (2026). Introducing Strands Decider 2B: a small, open source, decision model. Strands Agents blog, October 1, 2026. https://strandsagents.com/blog/introducing-strands-decider/ Verified 2026-10-04.
Perplexity (2026). Decisions API. Perplexity API documentation: pplx-decider-v1-27b, $0.04 per million input tokens, output free. https://docs.perplexity.ai/docs/decisions/quickstart Verified 2026-10-04.
Perplexity (2026). pplx-decider-v1-27b. Hugging Face model card: Qwen3.8-27B base, Apache 2.0, vendor benchmark table. https://huggingface.co/perplexity-ai/pplx-decider-v1-27b Verified 2026-10-04.
TypeSafe (2026). Models. TypeSafe documentation: jev-1.13.0 at $0.042 per million input tokens, output free. https://docs.typesafe.ai/models Verified for this site’s Jev post.
OpenAI (2026). API pricing. OpenAI developer documentation, GPT-6 Luna at $0.10 input and $0.50 output per million tokens. https://developers.openai.com/api/docs/pricing Verified 2026-09-23.
Unbiased (2026). Pareto 26.10 Preview model card. Unbiased, October 1, 2026: $0.80 input, $0.03 cached input, $3.20 output per million tokens; preliminary vendor-run benchmarks. https://unbiased.ai/model-card/ Verified 2026-10-04.
Microsoft AI (2026). Our first streaming transcription model debuts at no. 1 on Artificial Analysis. Microsoft AI news, October 1, 2026: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash, prices and availability. https://microsoft.ai/news/our-first-streaming-transcription-model/ Verified 2026-10-04.
Black Forest Labs (2026). Release notes. BFL documentation, FLUX 3 Image entry dated October 1, 2026, with per-image pricing by resolution. https://docs.bfl.ai/release-notes Verified 2026-10-04.
Tavus (2026). Griffin: The First Human Interaction Model. Tavus, October 1, 2026: Griffin-Lite research preview, VideoFDB scores, company-run human study. https://www.tavus.io/griffin Verified 2026-10-04.
Fastino Labs (2026). Introducing GLiDE: The First Thinking Decision Model. Fastino blog, September 30, 2026. https://fastino.ai/blog/introducing-glide-the-first-thinking-decision-model Verified 2026-10-04.
MarkTechPost (2026). Liquid AI Releases d1. Secondary coverage dating Liquid AI’s d1 to September 29, 2026. https://www.marktechpost.com/2026/09/29/liquid-ai-releases-d1-a-decision-model-that-returns-calibrated-probabilities-with-zero-output-tokens/ Read 2026-10-04.
Monthly release counts are derived from the release archive maintained on this site, whose every row carries the provider primary source it was verified against. The cost-per-1,000-decisions figures are this site’s arithmetic on the list prices above.