Best Open-Weight AI Models in 2026
The best open-weight AI models in 2026, ranked by use case: coding, long context, multimodal, on-device, and the real cost per finished task.
Leaderboards compress messy reality into a single number, then labs quote whichever number flatters the release. Suites saturate, contamination inflates scores, and the same model under different scaffolding can move ten points. What matters is which benchmarks still separate frontier models, and what each one refuses to measure.
This guide frames the system before you move into the newest signals.
The best open-weight AI models in 2026, ranked by use case: coding, long context, multimodal, on-device, and the real cost per finished task.
The featured guide stays above; the stream below moves as new analysis is published.
Muse Spark 1.2 at $1.25/$4.25 adds a $0.10/$0.20 contributor tier. Cost per task vs Claude Opus 5 and GPT-5.6 Sol, and what the data-for-discount really buys.
Terminal-Bench 4.0 cut to 66 tasks and Opus 5 leads at 51.8 percent. Derived cost per solved task runs from 6 to 234 dollars, a 38x spread.
GLM-5.3-Flash is Ox Alpha, officially: a 320B-A18B MoE model, 1M-token context, MIT license, and $0.15/$0.50 per million tokens. Specs, pricing, and benchmarks.
Ox Alpha is a free anonymous reasoning model on OpenRouter. Specs, the GLM-5.3 fingerprint case, benchmark rumors, and who pays for 100T tokens a day.
Apache 2.0 weights for Qwen3.8-27B landed on August 14, 2026. Strong vendor scores, hosted tokens at 80 percent off the flagship, and a catch in the local math.
Anthropic has published nothing on Fable 5.1. Four leaked claims, two X threads and one gray-scale sighting, graded against the primary sources.
Grok 4.6 kept the $2 and $6 headline but raised cached input 67 percent, and doubles every rate past 200K tokens. What that does to an agent bill.
AI writes 42 percent of committed code, yet the measured productivity gain stays contested. A sourced survey of the tools, techniques and evidence.
Every new AI model released in August 2026 with dates, verified per-token prices and primary sources, plus the price changes that landed alongside them.
Claude Opus 5 is stronger, but Sonnet 5 costs 40% less at standard rates. Compare coding, reasoning, speed, context, and when upgrading to Opus pays off.
Artificial Analysis scored Qwen3.8 Max at 58 on its Intelligence Index. The token price fell 20 percent but the cost to run the same eval rose 64 percent.
Prime Intellect reports 95.5% on ARC-AGI-3 with an open-source harness. The ARC-verified ceiling on that same set is 30.2%. What the gap measures.
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index at $0.14 per million input tokens, ahead of the pricier V4 Pro at 44.
Identical open weights score 67.8 or 90.0 on SWE-bench depending on the harness. What the research says about where local coding agents break.
Claude Opus 5 ships at $5/$25 per million tokens with an effort dial, matching Fable 5 intelligence at half the price. Benchmarks, specs, and cost per task.
A practitioner guide to LLM evaluation metrics: reference-based scores, LLM-as-judge, and human review, ranked by accuracy and cost per 10,000 responses.
Nearly 200 startups urged Trump not to restrict Chinese open-weight AI. Here is what a ban would target, why enforcement is hard, and the stakes.
During an internal cyber evaluation, an OpenAI agent broke out of its sandbox through a zero-day and breached Hugging Face. Here is what it means.
Qwen3.8 Max Preview is live through the Alibaba Token Plan and Qoder. Here is what is confirmed about access, benchmarks, open weights, and the 2.4T claim.
Chinese AI models already lead on price and open weights. A six-part scorecard shows why broad leadership is plausible by 2028–2030, but not inevitable.
Current 2026 API and subscription prices for DeepSeek, Qwen, Kimi, GLM, and MiniMax, compared per million tokens and against US models.
How AI cost per token fell roughly 10x per year since 2021, why it collapsed, and why your per-task bill did not drop nearly as fast.
Call Kimi K3 through Moonshot OpenAI-compatible endpoint or OpenRouter. Get a key, set the base URL, use model kimi-k3, and mind the max reasoning setting.
Kimi K3 and DeepSeek V4 are both open-weight Chinese models. K3 leads the AA intelligence index; DeepSeek V4 costs a fraction and leads on SWE-bench.
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
This topic collects 67 analyses, and they are written to be read together rather than one at a time: the featured guide sets out the shape of the problem, and the pieces below work through the individual numbers, tradeoffs and edge cases behind it. Every figure is attributed to a primary source at the point it is used and carries the date it was verified, because prices and benchmark results in this area go stale in weeks rather than years. Where a number is modeled rather than measured, the assumptions are stated so the arithmetic can be checked.
The topics on this site overlap by design. This one runs into Coding agents and Local AI & hardware, and a question that starts in one usually ends in another: a pricing decision turns into a hardware decision, a benchmark result turns into a cost question. Follow the links inside the posts rather than treating these archives as separate shelves.
Explore the value leaderboardTrack new and upcoming modelsLook up a benchmark