LMArena
Also known as Chatbot Arena, LMSYS Chatbot Arena, Arena
LMArena, previously LMSYS Chatbot Arena, ranks models by human preference. Users submit a prompt, see two anonymous responses, and vote for the better one. Those votes become a pairwise rating, similar to chess Elo. It answers "which model do people prefer talking to," which is a genuinely useful question and not the same question as "which model is more capable."
| What it measures | Crowdsourced human preference between two anonymized model responses, aggregated into a relative ranking, not an objective capability. |
|---|---|
| Built by | Arena (formerly LMArena and LMSYS Chatbot Arena; Angelopoulos, Chiang et al.), 2023 |
| Format | Open-ended head-to-head battles: users submit a prompt and vote on the better of two blind responses; tens of millions of votes |
| Scoring metric | Elo / Bradley-Terry pairwise rating (an Arena Score) |
| Status | Active |
| Representative top score | ~1510 Elo · Claude Opus 4.8 · read 2026-06 |
| Official leaderboard | arena.ai/leaderboard |
How LMArena works
A user prompt is sent to two randomly selected anonymised models. The user votes, and the result updates a Bradley-Terry style rating from which the Arena Score is derived. Aggregated over tens of millions of votes, the ranking is statistically stable. Because the prompts come from whatever real users happen to ask, the coverage is broad but uncontrolled, and there is no ground-truth answer anywhere in the process.
History and current status
The arena launched in 2023 and became the most-watched public ranking in AI, in large part because it resisted the contamination that was breaking static benchmarks: you cannot memorise a preference vote. It rebranded to Arena in early 2026. The same team now also runs Agent Arena, which scores agents from real usage using causal tracing rather than votes, in an explicit attempt to avoid the style-gaming problem.
What the score does not tell you
Preference is confounded with presentation. Response length, formatting, confident tone and markdown structure all raise vote share without raising correctness, so a model can climb by writing more attractively. The 2025 paper The Leaderboard Illusion argues further that private pre-release testing and uneven model deprecation bias ratings toward large labs with the resources to exploit both. Neither critique means the ranking is worthless; both mean it should not be read as a capability score.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
What LMArena scores actually mean
An Arena Score is an Elo-style rating from blind pairwise votes, so it measures which response people prefer, not which one is correct. That distinction governs everything about reading it. Style is a large confound: length, formatting, headers and confident tone all raise win rates independently of substance, which is why models tuned for presentation outrank models tuned for accuracy. Ratings are also only comparable within a snapshot, since the pool of competing models changes constantly. A 20-point gap between two models near the top is usually inside the confidence interval, and the leaderboard publishes those intervals precisely so that small gaps are not over-read.
Who reports LMArena, and how to read it
Labs cite Arena rank at launch, especially when their objective benchmark results are unremarkable. Read it as one axis of three: a preference rating, a composite index, and a hard unsaturated reasoning score. Where a launch leads with Arena rank alone, that is worth noticing. Rank differences inside the published confidence intervals should not be reported as a lead at all.
When to weight LMArena in a model choice
Use the Arena to answer one question well: which model will feel better to a general audience in open-ended chat. That is a real product question and no other benchmark here answers it. Do not use it to judge correctness, coding ability, tool use or factual reliability, all of which it explicitly does not measure. Weigh the critique in the 2025 paper The Leaderboard Illusion, which argues that private pre-release testing and uneven model deprecation bias ratings toward large labs, when interpreting rank differences. For capability, cross-check against LiveBench, which uses objective ground truth and no judges or votes.
Benchmarks to read alongside this one
Arena-Hard-Auto
Human-preference-aligned quality on hard open-ended prompts, scored automatically instead of by live human voting.
Copilot Arena
Which coding model developers actually prefer, collected from paired completions inside a real editor rather than a chat window.
AlpacaEval 2 (Length-Controlled)
Instruction-following quality judged by an LLM, with a regression correction for the judge’s bias toward longer answers.
LMArena: frequently asked questions
- What is LMArena?
- LMArena, formerly LMSYS Chatbot Arena, is a platform where users compare two anonymous model responses to their own prompt and vote for the better one. Those votes aggregate into a Bradley-Terry pairwise rating published as an Arena Score, based on tens of millions of votes.
- Is LMArena reliable?
- It reliably measures what it measures: human preference. It is not a capability measure. Ratings are influenced by response length, formatting and tone, and the 2025 paper The Leaderboard Illusion argues private testing and uneven deprecation bias ratings toward large labs.
- Can LMArena be gamed?
- Its ratings can be inflated without improving correctness, by optimising for the style humans vote for: longer answers, confident phrasing, heavy formatting. That is not cheating in the contamination sense, but it does mean rank movement can reflect presentation rather than substance.
Sources
- Chiang, Zheng, Sheng, Angelopoulos et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv preprint 2403.04132. arxiv.org/abs/2403.04132
- Arena (formerly LMArena). Official leaderboard with confidence intervals. arena.ai/leaderboard
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.