LMArena
Also known as Chatbot Arena, LMSYS Chatbot Arena, Arena
LMArena, previously LMSYS Chatbot Arena, ranks models by human preference. Users submit a prompt, see two anonymous responses, and vote for the better one. Those votes become a pairwise rating, similar to chess Elo. It answers "which model do people prefer talking to," which is a genuinely useful question and not the same question as "which model is more capable."
| What it measures | Crowdsourced human preference between two anonymized model responses, aggregated into a relative ranking, not an objective capability. |
|---|---|
| Built by | Arena (formerly LMArena and LMSYS Chatbot Arena; Angelopoulos, Chiang et al.), 2023 |
| Format | Open-ended head-to-head battles: users submit a prompt and vote on the better of two blind responses; tens of millions of votes |
| Scoring metric | Elo / Bradley-Terry pairwise rating (an Arena Score) |
| Status | Active |
| Representative top score | ~1510 Elo · Claude Opus 4.8 · read 2026-06 |
| Official leaderboard | arena.ai/leaderboard |
How LMArena works
A user prompt is sent to two randomly selected anonymised models. The user votes, and the result updates a Bradley-Terry style rating from which the Arena Score is derived. Aggregated over tens of millions of votes, the ranking is statistically stable. Because the prompts come from whatever real users happen to ask, the coverage is broad but uncontrolled, and there is no ground-truth answer anywhere in the process.
History and current status
The arena launched in 2023 and became the most-watched public ranking in AI, in large part because it resisted the contamination that was breaking static benchmarks: you cannot memorise a preference vote. It rebranded to Arena in early 2026. The same team now also runs Agent Arena, which scores agents from real usage using causal tracing rather than votes, in an explicit attempt to avoid the style-gaming problem.
What the score does not tell you
Preference is confounded with presentation. Response length, formatting, confident tone and markdown structure all raise vote share without raising correctness, so a model can climb by writing more attractively. The 2025 paper The Leaderboard Illusion argues further that private pre-release testing and uneven model deprecation bias ratings toward large labs with the resources to exploit both. Neither critique means the ranking is worthless; both mean it should not be read as a capability score.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports LMArena, and how to read it
Labs cite Arena rank at launch, especially when their objective benchmark results are unremarkable. Read it as one axis of three: a preference rating, a composite index, and a hard unsaturated reasoning score. Where a launch leads with Arena rank alone, that is worth noticing.
Benchmarks to read alongside this one
Arena-Hard-Auto
Human-preference-aligned quality on hard open-ended prompts, scored automatically instead of by live human voting.
Copilot Arena
Which coding model developers actually prefer, collected from paired completions inside a real editor rather than a chat window.
AlpacaEval 2 (Length-Controlled)
Instruction-following quality judged by an LLM, with a regression correction for the judge’s bias toward longer answers.
LMArena: frequently asked questions
- What is LMArena?
- LMArena, formerly LMSYS Chatbot Arena, is a platform where users compare two anonymous model responses to their own prompt and vote for the better one. Those votes aggregate into a Bradley-Terry pairwise rating published as an Arena Score, based on tens of millions of votes.
- Is LMArena reliable?
- It reliably measures what it measures: human preference. It is not a capability measure. Ratings are influenced by response length, formatting and tone, and the 2025 paper The Leaderboard Illusion argues private testing and uneven deprecation bias ratings toward large labs.
- Can LMArena be gamed?
- Its ratings can be inflated without improving correctness, by optimising for the style humans vote for: longer answers, confident phrasing, heavy formatting. That is not cheating in the contamination sense, but it does mean rank movement can reflect presentation rather than substance.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.