Capital & Compute

LegalBench

Benchmark· Domain & professional work· Checked 2026-07-27

LegalBench measures legal reasoning through 162 tasks designed by legal professionals rather than by machine-learning researchers. It is organised around the kinds of reasoning lawyers actually do, which makes the per-task results far more useful than any single aggregate score for deciding whether a model can be trusted with a specific legal workflow.

Key facts about the LegalBench benchmark
What it measuresLegal reasoning across the specific skills lawyers actually use, as defined by legal professionals.
Built byGuha, Nyarko, Ho et al., 2023
Format162 tasks hand-built by legal practitioners, covering issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding
Scoring metricPer-task accuracy, aggregated by reasoning type
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard
Official leaderboardhazyresearch.stanford.edu/legalbench

How LegalBench works

The tasks cover six types of legal reasoning: issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding. Each was hand-built by practitioners so that it either measures something practically useful or something lawyers consider intellectually central. Results are reported per task and aggregated by reasoning type, which is the level at which they are actionable.

History and current status

The benchmark was introduced in a 2023 paper with 40 authors, many of them practising lawyers, through a collaborative interdisciplinary process. It remains the reference legal-reasoning benchmark. LegalBench-RAG later extended the design to the retrieval half of legal question answering, which is where most production legal AI failures actually occur.

What the score does not tell you

A benchmark of discrete reasoning tasks does not capture what legal work consists of: long documents, conflicting authority, jurisdictional variation and consequences for being wrong. Aggregating 162 heterogeneous tasks into one number is close to meaningless, and per-task variance is wide. Coverage also skews toward United States law, so a strong score says little about another jurisdiction.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

LegalBench is not one of the 15 benchmarks scored in the benchmark trust ranking, which rates the most-quoted tests on those same four axes and sorts them most trustworthy first.

What LegalBench scores actually mean

LegalBench deliberately resists a single headline number. It is 162 separate tasks grouped by reasoning type, and an average across them mixes issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding into one figure that describes nothing in particular. Rule recall rewards memorisation and scores high; rule application and interpretation demand reasoning over a fact pattern and score much lower. A model can look strong on the aggregate while failing the reasoning types that matter for any real legal workflow, which is why the per-task and per-category tables are the substance of the benchmark and the mean is close to noise.

Who reports LegalBench, and how to read it

LegalBench is cited in academic work and by legal-technology vendors, and it appears less often in frontier model cards than medical benchmarks do. For evaluating a legal AI product, the useful move is to read the specific tasks that match your workflow rather than any headline figure. A single averaged LegalBench score is uninformative by construction, so treat one as a red flag.

When to weight LegalBench in a model choice

Use LegalBench by picking the reasoning types that match the workflow, then reading those tasks alone. Contract review draws on interpretation and rule application; research assistance draws more on recall. Ignore any vendor claim of a single LegalBench score, because the construction of the benchmark makes such a number uninformative by design. Legal work is high-stakes and jurisdiction-specific, so treat even strong per-task results as a screening tool for which models to evaluate on your own documents rather than as evidence of fitness for use.

2023
First released
Guha, Nyarko, Ho et al.
Active
Status today
As of September 4, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

LegalBench: frequently asked questions

What is LegalBench?
LegalBench is a 2023 benchmark of 162 legal reasoning tasks hand-built by legal professionals, covering issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding. Results are reported per task and by reasoning type.
Does a good LegalBench score mean a model can do legal work?
No. The tasks are discrete reasoning problems, while legal work involves long documents, conflicting authority, jurisdictional variation and real consequences. Read the specific tasks that match your workflow rather than the aggregate, and note the United States law skew.
What is LegalBench-RAG?
LegalBench-RAG extends the benchmark to the retrieval side of legal question answering, evaluating whether the right passages are found at all. That matters because most production failures in legal AI come from retrieval rather than from reasoning over correctly retrieved text.

Sources

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directoryWhich benchmarks are worth trusting →