Capital & Compute

LMArena

Benchmark· Human preference & holistic· Checked 2026-07-03

Also known as Chatbot Arena, LMSYS Chatbot Arena, Arena

LMArena, previously LMSYS Chatbot Arena, ranks models by human preference. Users submit a prompt, see two anonymous responses, and vote for the better one. Those votes become a pairwise rating, similar to chess Elo. It answers "which model do people prefer talking to," which is a genuinely useful question and not the same question as "which model is more capable."

Key facts about the LMArena benchmark
What it measuresCrowdsourced human preference between two anonymized model responses, aggregated into a relative ranking, not an objective capability.
Built byArena (formerly LMArena and LMSYS Chatbot Arena; Angelopoulos, Chiang et al.), 2023
FormatOpen-ended head-to-head battles: users submit a prompt and vote on the better of two blind responses; tens of millions of votes
Scoring metricElo / Bradley-Terry pairwise rating (an Arena Score)
StatusActive
Representative top score~1510 Elo · Claude Opus 4.8 · read 2026-06
Official leaderboardarena.ai/leaderboard

How LMArena works

A user prompt is sent to two randomly selected anonymised models. The user votes, and the result updates a Bradley-Terry style rating from which the Arena Score is derived. Aggregated over tens of millions of votes, the ranking is statistically stable. Because the prompts come from whatever real users happen to ask, the coverage is broad but uncontrolled, and there is no ground-truth answer anywhere in the process.

History and current status

The arena launched in 2023 and became the most-watched public ranking in AI, in large part because it resisted the contamination that was breaking static benchmarks: you cannot memorise a preference vote. It rebranded to Arena in early 2026. The same team now also runs Agent Arena, which scores agents from real usage using causal tracing rather than votes, in an explicit attempt to avoid the style-gaming problem.

What the score does not tell you

Preference is confounded with presentation. Response length, formatting, confident tone and markdown structure all raise vote share without raising correctness, so a model can climb by writing more attractively. The 2025 paper The Leaderboard Illusion argues further that private pre-release testing and uneven model deprecation bias ratings toward large labs with the resources to exploit both. Neither critique means the ranking is worthless; both mean it should not be read as a capability score.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

What LMArena scores actually mean

An Arena Score is an Elo-style rating from blind pairwise votes, so it measures which response people prefer, not which one is correct. That distinction governs everything about reading it. Style is a large confound: length, formatting, headers and confident tone all raise win rates independently of substance, which is why models tuned for presentation outrank models tuned for accuracy. Ratings are also only comparable within a snapshot, since the pool of competing models changes constantly. A 20-point gap between two models near the top is usually inside the confidence interval, and the leaderboard publishes those intervals precisely so that small gaps are not over-read.

Who reports LMArena, and how to read it

Labs cite Arena rank at launch, especially when their objective benchmark results are unremarkable. Read it as one axis of three: a preference rating, a composite index, and a hard unsaturated reasoning score. Where a launch leads with Arena rank alone, that is worth noticing. Rank differences inside the published confidence intervals should not be reported as a lead at all.

When to weight LMArena in a model choice

Use the Arena to answer one question well: which model will feel better to a general audience in open-ended chat. That is a real product question and no other benchmark here answers it. Do not use it to judge correctness, coding ability, tool use or factual reliability, all of which it explicitly does not measure. Weigh the critique in the 2025 paper The Leaderboard Illusion, which argues that private pre-release testing and uneven model deprecation bias ratings toward large labs, when interpreting rank differences. For capability, cross-check against LiveBench, which uses objective ground truth and no judges or votes.

2023
First released
Arena (formerly LMArena and LMSYS Chatbot Arena; Angelopoulos, Chiang et al.)
Active
Status today
As of September 4, 2026
~1510 Elo
Representative top score
Read 2026-06

Benchmarks to read alongside this one

LMArena: frequently asked questions

What is LMArena?
LMArena, formerly LMSYS Chatbot Arena, is a platform where users compare two anonymous model responses to their own prompt and vote for the better one. Those votes aggregate into a Bradley-Terry pairwise rating published as an Arena Score, based on tens of millions of votes.
Is LMArena reliable?
It reliably measures what it measures: human preference. It is not a capability measure. Ratings are influenced by response length, formatting and tone, and the 2025 paper The Leaderboard Illusion argues private testing and uneven deprecation bias ratings toward large labs.
Can LMArena be gamed?
Its ratings can be inflated without improving correctness, by optimising for the style humans vote for: longer answers, confident phrasing, heavy formatting. That is not cheating in the contamination sense, but it does mean rank movement can reflect presentation rather than substance.

Sources

  • Chiang, Zheng, Sheng, Angelopoulos et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv preprint 2403.04132. arxiv.org/abs/2403.04132
  • Arena (formerly LMArena). Official leaderboard with confidence intervals. arena.ai/leaderboard

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory