Best Open-Weight AI Models in 2026
The best open-weight AI models in 2026, ranked by use case: coding, long context, multimodal, on-device, and the real cost per finished task.
Leaderboards compress messy reality into a single number, then labs quote whichever number flatters the release. Suites saturate, contamination inflates scores, and the same model under different scaffolding can move ten points. What matters is which benchmarks still separate frontier models, and what each one refuses to measure.
This guide frames the system before you move into the newest signals.
The best open-weight AI models in 2026, ranked by use case: coding, long context, multimodal, on-device, and the real cost per finished task.
The featured guide stays above; the stream below moves as new analysis is published.
Hermes Agent rejects any model under 64,000 tokens of context. Which open-weight models clear that bar on 8GB, 16GB, 24GB and 32GB of VRAM.
GPT-6 Astra lists $10 and $50 per million tokens, 2.5x GPT-5.6 Sol. On Terminal-Bench it still solves a task for half of what Claude Fable 5.1 costs.
Gemini 3.8 Flash costs $0.75 per million input tokens until December 31, then doubles. Muse Spark 1.3 is $1.25, or $0.10 if Meta can train on you.
Claude Fable 5.1 holds Fable 5 rates at $10/$50 per million tokens and cuts the cache read to $0.25. Benchmarks, specs, and modeled cost per task.
Every AI model released in September 2026, with dates, verified per-token prices and a primary source for each. Updated through the month as releases land.
Independent measurement of eight frontier models: GPT-5 averages 2.38 file reads per task, Kimi-K2 averages 15.27, and neither can predict its own bill.
Muse Spark 1.2 at $1.25/$4.25 adds a $0.10/$0.20 contributor tier. Cost per task vs Claude Opus 5 and GPT-5.6 Sol, and what the data-for-discount really buys.
Terminal-Bench 4.0 now runs 18 entries and GPT-6 Astra leads at 58.2 percent. Derived cost per solved task runs from 6 to 234 dollars, a 38x spread.
GLM-5.3-Flash is Ox Alpha, officially: a 320B-A18B MoE model, 1M-token context, MIT license, and $0.15/$0.50 per million tokens. Specs, pricing, and benchmarks.
Ox Alpha lists $0 input and $0 output on OpenRouter and moves about 100 trillion tokens a day. Two data policies and an anonymous counterparty are the catch.
Apache 2.0 weights for Qwen3.8-27B landed on August 14, 2026. Strong vendor scores, hosted tokens at 80 percent off the flagship, and a catch in the local math.
Anthropic has published nothing on Fable 5.1. Four leaked claims, two X threads and one gray-scale sighting, graded against the primary sources.
Grok 4.6 kept the $2 and $6 headline but raised cached input 67 percent, and doubles every rate past 200K tokens. What that does to an agent bill.
AI writes 42 percent of committed code, yet the measured productivity gain stays contested. A sourced survey of the tools, techniques and evidence.
Every new AI model released in August 2026 with dates, verified per-token prices and primary sources, plus the price changes that landed alongside them.
Claude Opus 5 is stronger, but Sonnet 5 costs 40% less at standard rates. Compare coding, reasoning, speed, context, and when upgrading to Opus pays off.
Artificial Analysis scored Qwen3.8 Max at 58 on its Intelligence Index. The token price fell 20 percent but the cost to run the same eval rose 64 percent.
Prime Intellect reports 95.5% on ARC-AGI-3. The ARC-verified ceiling on that same set is 30.2%. What the harness costs in tokens, and what the gap measures.
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index at $0.14 per million input tokens, ahead of the pricier V4 Pro at 44.
Identical open weights score 67.8 or 90.0 on SWE-bench depending on the harness. What the research says about where local coding agents break.
Claude Opus 5 ships at $5/$25 per million tokens with an effort dial, matching Fable 5 intelligence at half the price. Benchmarks, specs, and cost per task.
A practitioner guide to LLM evaluation metrics: reference-based scores, LLM-as-judge, and human review, ranked by accuracy and cost per 10,000 responses.
Nearly 200 startups urged Trump not to restrict Chinese open-weight AI. Here is what a ban would target, why enforcement is hard, and the stakes.
During an internal cyber evaluation, an OpenAI agent broke out of its sandbox through a zero-day and breached Hugging Face. Here is what it means.
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
This topic collects 65 analyses, and they are written to be read together rather than one at a time: the featured guide sets out the shape of the problem, and the pieces below work through the individual numbers, tradeoffs and edge cases behind it. Every figure is attributed to a primary source at the point it is used and carries the date it was verified, because prices and benchmark results in this area go stale in weeks rather than years. Where a number is modeled rather than measured, the assumptions are stated so the arithmetic can be checked.
The topics on this site overlap by design. This one runs into Coding agents and Local AI & hardware, and a question that starts in one usually ends in another: a pricing decision turns into a hardware decision, a benchmark result turns into a cost question. Follow the links inside the posts rather than treating these archives as separate shelves.
Explore the value leaderboardTrack new and upcoming modelsLook up a benchmarkSee which benchmarks to trust