Capital & Compute

LiveBench

Benchmark· Human preference & holistic· Checked 2026-07-27

LiveBench is a composite benchmark that scores six capability categories at once and refreshes its questions every month. It has no LLM judge and no human voting: every question has a verifiable ground-truth answer. That combination makes it the strongest available alternative to preference arenas for a single broad capability number.

Key facts about the LiveBench benchmark
What it measuresBroad capability across six categories at once (math, coding, reasoning, data analysis, instruction following, language), on questions refreshed monthly.
Built byWhite, Dooley, Roberts et al., 2024
FormatQuestions drawn from recent math competitions, arXiv papers, news and datasets, with roughly one sixth replaced each month so the set fully refreshes about every six months
Scoring metricObjective automatic scoring against ground truth, averaged across categories
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard
Official leaderboardlivebench.ai

How LiveBench works

The six categories are mathematics, coding, reasoning, data analysis, instruction following and language comprehension. Questions are drawn from recent sources such as new math competitions, arXiv papers, news articles and datasets, so they postdate most training cutoffs. Roughly one sixth of the questions are replaced each month, meaning the set fully refreshes about every six months. Scoring is automatic against ground truth.

History and current status

A large author group including Colin White, Samuel Dooley and Yann LeCun introduced it in 2024, and it was accepted as a Spotlight at ICLR 2025. It has been maintained on its monthly refresh cadence since. Notably, the authors renamed it between versions from "A Challenging, Contamination-Free LLM Benchmark" to "Contamination-Limited," which is a more honest description of what monthly rotation can achieve.

What the score does not tell you

Monthly rotation limits contamination but does not eliminate it, as the authors’ own retitling concedes. The rotation has a side effect: scores from different months are not strictly comparable, because the questions differ, so tracking a model over time on LiveBench requires care. Objective ground-truth scoring also restricts it to questions with checkable answers, which excludes most open-ended work.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Scored on those four axes, LiveBench carries a concern score of 1 out of 8 and ranks 3 of 15, which puts it in the group worth quoting as it stands. See the reasoning behind that rating and how it compares with the rest of the field.

What LiveBench scores actually mean

LiveBench scores are averages across categories drawn from recent competitions, arXiv papers, news and datasets, with roughly one sixth of questions replaced monthly so the set fully refreshes about every six months. That rotation is what keeps the numbers meaningful, and it also means a score is tied to a snapshot: comparing a figure from one month against another is comparing two partly different benchmarks. The category breakdown matters more than the mean, since a model can be strong on math and weak on language reasoning and land mid-table. Because scoring is objective and automatic with no LLM judge and no human votes, the results avoid the style bias that dominates preference arenas.

Who reports LiveBench, and how to read it

LiveBench is maintained independently and publishes its own leaderboard. It appears in academic comparisons more than in launch marketing, partly because a rotating benchmark is inconvenient for a fixed announcement. That independence is precisely why it is worth reading alongside a lab’s own numbers. Record the snapshot month with any figure, since monthly rotation makes an undated LiveBench number unreproducible.

When to weight LiveBench in a model choice

Use LiveBench as the objective counterweight to the Arena. When a model ranks high on human preference and mid-table on LiveBench, the gap is usually presentation rather than capability, and that is a useful thing to know before choosing a model for analytical work. Read the per-category columns for the categories you care about. Record the snapshot date with any figure you quote, because monthly rotation makes undated LiveBench numbers unreproducible. Note the authors' own framing: they renamed the set from contamination-free to contamination-limited, which is the accurate description.

2024
First released
White, Dooley, Roberts et al.
Active
Status today
As of September 12, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

LiveBench: frequently asked questions

What is LiveBench?
LiveBench is a composite LLM benchmark covering mathematics, coding, reasoning, data analysis, instruction following and language comprehension. Questions come from recent sources and roughly one sixth are replaced monthly, so the set fully refreshes about every six months.
Why is LiveBench considered contamination-resistant?
Because its questions are drawn from recent sources that postdate most training cutoffs and are rotated monthly, so memorising them has limited value. The authors describe it as contamination-limited rather than contamination-free, which is the accurate framing.
Does LiveBench use an LLM judge?
No. Every question has a verifiable objective ground-truth answer and is scored automatically. That avoids both the bias of LLM-as-judge scoring and the style effects of human preference voting.

Sources

  • White, Dooley, Roberts, Pal et al. (2024). LiveBench: A Challenging, Contamination-Limited LLM Benchmark. arXiv preprint 2406.19314. arxiv.org/abs/2406.19314
  • LiveBench. Official leaderboard with per-category results and refresh dates. livebench.ai

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →