Capital & Compute
Benchmark· Human preference & holistic· Checked 2026-07-27

LiveBench

LiveBench is a composite benchmark that scores six capability categories at once and refreshes its questions every month. It has no LLM judge and no human voting: every question has a verifiable ground-truth answer. That combination makes it the strongest available alternative to preference arenas for a single broad capability number.

Key facts about the LiveBench benchmark
What it measuresBroad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly.
Built byWhite, Dooley, Roberts et al., 2024
FormatQuestions drawn from recent math competitions, arXiv papers, news and datasets, with roughly one sixth replaced each month so the set fully refreshes about every six months
Scoring metricObjective automatic scoring against ground truth, averaged across categories
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard
Official leaderboardlivebench.ai

How LiveBench works

The six categories are mathematics, coding, reasoning, data analysis, instruction following and language comprehension. Questions are drawn from recent sources such as new math competitions, arXiv papers, news articles and datasets, so they postdate most training cutoffs. Roughly one sixth of the questions are replaced each month, meaning the set fully refreshes about every six months. Scoring is automatic against ground truth.

History and current status

A large author group including Colin White, Samuel Dooley and Yann LeCun introduced it in 2024, and it was accepted as a Spotlight at ICLR 2025. It has been maintained on its monthly refresh cadence since. Notably, the authors renamed it between versions from "A Challenging, Contamination-Free LLM Benchmark" to "Contamination-Limited," which is a more honest description of what monthly rotation can achieve.

What the score does not tell you

Monthly rotation limits contamination but does not eliminate it, as the authors’ own retitling concedes. The rotation has a side effect: scores from different months are not strictly comparable, because the questions differ, so tracking a model over time on LiveBench requires care. Objective ground-truth scoring also restricts it to questions with checkable answers, which excludes most open-ended work.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports LiveBench, and how to read it

LiveBench is maintained independently and publishes its own leaderboard. It appears in academic comparisons more than in launch marketing, partly because a rotating benchmark is inconvenient for a fixed announcement. That independence is precisely why it is worth reading alongside a lab’s own numbers.

2024
First released
White, Dooley, Roberts et al.
Active
Status today
As of July 27, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

LiveBench: frequently asked questions

What is LiveBench?
LiveBench is a composite LLM benchmark covering mathematics, coding, reasoning, data analysis, instruction following and language comprehension. Questions come from recent sources and roughly one sixth are replaced monthly, so the set fully refreshes about every six months.
Why is LiveBench considered contamination-resistant?
Because its questions are drawn from recent sources that postdate most training cutoffs and are rotated monthly, so memorising them has limited value. The authors describe it as contamination-limited rather than contamination-free, which is the accurate framing.
Does LiveBench use an LLM judge?
No. Every question has a verifiable objective ground-truth answer and is scored automatically. That avoids both the bias of LLM-as-judge scoring and the style effects of human preference voting.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory