Hermes Agent Explained: Why the Harness Beats the Model
What Hermes Agent is in 2026: an MIT-licensed agent with a learning loop, 300-plus models and a v0.21.0 provider wave. Why the harness beats the model.
Leaderboards compress messy reality into a single number, then labs quote whichever number flatters the release. Suites saturate, contamination inflates scores, and the same model under different scaffolding can move ten points. What matters is which benchmarks still separate frontier models, and what each one refuses to measure.
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
What Hermes Agent is in 2026: an MIT-licensed agent with a learning loop, 300-plus models and a v0.21.0 provider wave. Why the harness beats the model.
OpenAI is shutting Sora down while Chinese models top the independent video leaderboards on quality and cost. The state of AI video, July 2026.
Arena, formerly LMArena, ranks AI agents from over a million real sessions using causal tracing, not style votes. How Agent Arena works and who leads.
Chinese AI models list output tokens up to 57x below US flagships. The verified economics of efficient training, cheap power and open weights as strategy.
Benchmark saturation is when top AI models bunch so close to the ceiling a test cannot rank them. Why MMLU hit this wall, and the harder replacement is next.
Claude Fable 5 returns after a US export-control suspension. Here is the $10/$50 pricing, the SWE-Bench Pro and FrontierCode scores, and what it means.
Claude Sonnet 5 ships at $3/$15 per million tokens, intro $2/$10 through August 2026. What changed versus Sonnet 4.6 and Opus 4.8, and the cost per task.
DeepSWE grades whether an AI finishes the task. FrontierCode grades whether you would merge its code. Why the same model scores 59% and 13%.
OpenAI cut GPT-5.6 API prices on July 30: Luna by 80 percent, Terra by 20, Sol not at all. On August 21 the flagship followed with a promotional cut to $4/$20.
OpenAI launched GPT-5.6 on June 26 as three tiers: Sol, Terra and Luna. Here are the confirmed prices, the benchmarks, and the government access catch.
The Trump administration asked OpenAI to stagger GPT-5.6 and approve users one by one. A federal gate now sits between a finished AI model and the public.
How much RAM a local LLM really needs: measured Q4_K_M file sizes, what 8GB to 512GB runs in 2026, and why long context can cost more than the weights.
Benchmark contamination is when test answers leak into training data and inflate AI scores. How it happens, how much it distorts results, and how to spot it.
A leaked gpt-5.6-preview route lit up r/OpenAI, then GPT-5.6 launched June 26 as Sol, Terra and Luna in a government-gated preview. The signal vs the hype.
AI benchmark scores get gamed by contamination, saturation, and cheating. The 2026 receipts, the four failure modes, and how to read past any leaderboard.
Claude Fable 5 costs exactly 2x Opus 4.8 per token: $10/$50 vs $5/$25. Whether it is cheaper per task depends on loop count, not the sticker rate.
Codex leads Claude Code 83.4% to 78.9% on Terminal-Bench 2.1, but the prices match and the cheaper agent flips with your harness. The honest scorecard.
A 2026 Microsoft Research preprint found the cheaper-per-token AI model cost more to finish the job in 32% of model pairs. Why the sticker misleads.
GPT-5.6 Sol lists $4 input and $20 output per million tokens, promotional through November 21 2026. What that works out to per finished task.
Token data from OpenRouter shows open-source LLMs passing proprietary models in mid-2026, a roughly 60/40 flip. The daily breakdown by AI lab.
Open-weights LLMs crossed from toy to useful in 2026. What actually changed, and the cost math for when running a model yourself beats paying an API.
AI agent benchmarks broke in 2026: reward-hacked to 100%, SWE-bench Verified retired, scores swung by harness choice. What each measures and which to trust.
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
This topic collects 71 analyses, and they are written to be read together rather than one at a time: the featured guide sets out the shape of the problem, and the pieces below work through the individual numbers, tradeoffs and edge cases behind it. Every figure is attributed to a primary source at the point it is used and carries the date it was verified, because prices and benchmark results in this area go stale in weeks rather than years. Where a number is modeled rather than measured, the assumptions are stated so the arithmetic can be checked.
The topics on this site overlap by design. This one runs into Coding agents and Local AI & hardware, and a question that starts in one usually ends in another: a pricing decision turns into a hardware decision, a benchmark result turns into a cost question. Follow the links inside the posts rather than treating these archives as separate shelves.
Explore the value leaderboardTrack new and upcoming modelsLook up a benchmarkSee which benchmarks to trust