LiveBench
LiveBench is a composite benchmark that scores six capability categories at once and refreshes its questions every month. It has no LLM judge and no human voting: every question has a verifiable ground-truth answer. That combination makes it the strongest available alternative to preference arenas for a single broad capability number.
| What it measures | Broad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly. |
|---|---|
| Built by | White, Dooley, Roberts et al., 2024 |
| Format | Questions drawn from recent math competitions, arXiv papers, news and datasets, with roughly one sixth replaced each month so the set fully refreshes about every six months |
| Scoring metric | Objective automatic scoring against ground truth, averaged across categories |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
| Official leaderboard | livebench.ai |
How LiveBench works
The six categories are mathematics, coding, reasoning, data analysis, instruction following and language comprehension. Questions are drawn from recent sources such as new math competitions, arXiv papers, news articles and datasets, so they postdate most training cutoffs. Roughly one sixth of the questions are replaced each month, meaning the set fully refreshes about every six months. Scoring is automatic against ground truth.
History and current status
A large author group including Colin White, Samuel Dooley and Yann LeCun introduced it in 2024, and it was accepted as a Spotlight at ICLR 2025. It has been maintained on its monthly refresh cadence since. Notably, the authors renamed it between versions from "A Challenging, Contamination-Free LLM Benchmark" to "Contamination-Limited," which is a more honest description of what monthly rotation can achieve.
What the score does not tell you
Monthly rotation limits contamination but does not eliminate it, as the authors’ own retitling concedes. The rotation has a side effect: scores from different months are not strictly comparable, because the questions differ, so tracking a model over time on LiveBench requires care. Objective ground-truth scoring also restricts it to questions with checkable answers, which excludes most open-ended work.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports LiveBench, and how to read it
LiveBench is maintained independently and publishes its own leaderboard. It appears in academic comparisons more than in launch marketing, partly because a rotating benchmark is inconvenient for a fixed announcement. That independence is precisely why it is worth reading alongside a lab’s own numbers.
Benchmarks to read alongside this one
LiveCodeBench
Code generation and related skills (self-repair, execution, test-output prediction) on fresh competitive-programming problems, designed to be contamination-free.
Epoch Capabilities Index
Overall model capability on one continuous scale, stitched together from many benchmarks of differing difficulty.
Arena-Hard-Auto
Human-preference-aligned quality on hard open-ended prompts, scored automatically instead of by live human voting.
LiveBench: frequently asked questions
- What is LiveBench?
- LiveBench is a composite LLM benchmark covering mathematics, coding, reasoning, data analysis, instruction following and language comprehension. Questions come from recent sources and roughly one sixth are replaced monthly, so the set fully refreshes about every six months.
- Why is LiveBench considered contamination-resistant?
- Because its questions are drawn from recent sources that postdate most training cutoffs and are rotated monthly, so memorising them has limited value. The authors describe it as contamination-limited rather than contamination-free, which is the accurate framing.
- Does LiveBench use an LLM judge?
- No. Every question has a verifiable objective ground-truth answer and is scored automatically. That avoids both the bias of LLM-as-judge scoring and the style effects of human preference voting.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.