Capital & Compute
Benchmark· Domain & professional work· Checked 2026-07-27

LegalBench

LegalBench measures legal reasoning through 162 tasks designed by legal professionals rather than by machine-learning researchers. It is organised around the kinds of reasoning lawyers actually do, which makes the per-task results far more useful than any single aggregate score for deciding whether a model can be trusted with a specific legal workflow.

Key facts about the LegalBench benchmark
What it measuresLegal reasoning across the specific skills lawyers actually use, as defined by legal professionals.
Built byGuha, Nyarko, Ho et al., 2023
Format162 tasks hand-built by legal practitioners, covering issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding
Scoring metricPer-task accuracy, aggregated by reasoning type
StatusActive
Representative top scoreNot independently confirmed; see the leaderboard
Official leaderboardhazyresearch.stanford.edu/legalbench

How LegalBench works

The tasks cover six types of legal reasoning: issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding. Each was hand-built by practitioners so that it either measures something practically useful or something lawyers consider intellectually central. Results are reported per task and aggregated by reasoning type, which is the level at which they are actionable.

History and current status

The benchmark was introduced in a 2023 paper with 40 authors, many of them practising lawyers, through a collaborative interdisciplinary process. It remains the reference legal-reasoning benchmark. LegalBench-RAG later extended the design to the retrieval half of legal question answering, which is where most production legal AI failures actually occur.

What the score does not tell you

A benchmark of discrete reasoning tasks does not capture what legal work consists of: long documents, conflicting authority, jurisdictional variation and consequences for being wrong. Aggregating 162 heterogeneous tasks into one number is close to meaningless, and per-task variance is wide. Coverage also skews toward United States law, so a strong score says little about another jurisdiction.

This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.

Who reports LegalBench, and how to read it

LegalBench is cited in academic work and by legal-technology vendors, and it appears less often in frontier model cards than medical benchmarks do. For evaluating a legal AI product, the useful move is to read the specific tasks that match your workflow rather than any headline figure.

2023
First released
Guha, Nyarko, Ho et al.
Active
Status today
As of July 27, 2026
See board
Representative top score
Not independently confirmed

Benchmarks to read alongside this one

LegalBench: frequently asked questions

What is LegalBench?
LegalBench is a 2023 benchmark of 162 legal reasoning tasks hand-built by legal professionals, covering issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding. Results are reported per task and by reasoning type.
Does a good LegalBench score mean a model can do legal work?
No. The tasks are discrete reasoning problems, while legal work involves long documents, conflicting authority, jurisdictional variation and real consequences. Read the specific tasks that match your workflow rather than the aggregate, and note the United States law skew.
What is LegalBench-RAG?
LegalBench-RAG extends the benchmark to the retrieval side of legal question answering, evaluating whether the right passages are found at all. That matters because most production failures in legal AI come from retrieval rather than from reasoning over correctly retrieved text.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← All 116 benchmarks in the directory