LegalBench
LegalBench measures legal reasoning through 162 tasks designed by legal professionals rather than by machine-learning researchers. It is organised around the kinds of reasoning lawyers actually do, which makes the per-task results far more useful than any single aggregate score for deciding whether a model can be trusted with a specific legal workflow.
| What it measures | Legal reasoning across the specific skills lawyers actually use, as defined by legal professionals. |
|---|---|
| Built by | Guha, Nyarko, Ho et al., 2023 |
| Format | 162 tasks hand-built by legal practitioners, covering issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding |
| Scoring metric | Per-task accuracy, aggregated by reasoning type |
| Status | Active |
| Representative top score | Not independently confirmed; see the leaderboard |
| Official leaderboard | hazyresearch.stanford.edu/legalbench |
How LegalBench works
The tasks cover six types of legal reasoning: issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding. Each was hand-built by practitioners so that it either measures something practically useful or something lawyers consider intellectually central. Results are reported per task and aggregated by reasoning type, which is the level at which they are actionable.
History and current status
The benchmark was introduced in a 2023 paper with 40 authors, many of them practising lawyers, through a collaborative interdisciplinary process. It remains the reference legal-reasoning benchmark. LegalBench-RAG later extended the design to the retrieval half of legal question answering, which is where most production legal AI failures actually occur.
What the score does not tell you
A benchmark of discrete reasoning tasks does not capture what legal work consists of: long documents, conflicting authority, jurisdictional variation and consequences for being wrong. Aggregating 162 heterogeneous tasks into one number is close to meaningless, and per-task variance is wide. Coverage also skews toward United States law, so a strong score says little about another jurisdiction.
This is the general case, not a quirk of one benchmark. For the four failure modes that decide whether any benchmark number means anything, read are AI benchmarks reliable, and for this row in particular, how benchmark contamination inflates scores.
Who reports LegalBench, and how to read it
LegalBench is cited in academic work and by legal-technology vendors, and it appears less often in frontier model cards than medical benchmarks do. For evaluating a legal AI product, the useful move is to read the specific tasks that match your workflow rather than any headline figure.
Benchmarks to read alongside this one
FinanceBench
Open-book financial question answering over real public-company filings, with the supporting evidence required.
HealthBench
Open-ended clinical conversation quality and safety, graded against rubrics written by practising physicians.
MedHELM
Clinical ability across the breadth of real medical work, on a clinician-validated taxonomy rather than exam questions.
LegalBench: frequently asked questions
- What is LegalBench?
- LegalBench is a 2023 benchmark of 162 legal reasoning tasks hand-built by legal professionals, covering issue spotting, rule recall, rule application, rule conclusion, interpretation and rhetorical understanding. Results are reported per task and by reasoning type.
- Does a good LegalBench score mean a model can do legal work?
- No. The tasks are discrete reasoning problems, while legal work involves long documents, conflicting authority, jurisdictional variation and real consequences. Read the specific tasks that match your workflow rather than the aggregate, and note the United States law skew.
- What is LegalBench-RAG?
- LegalBench-RAG extends the benchmark to the retrieval side of legal question answering, evaluating whether the right passages are found at all. That matters because most production failures in legal AI come from retrieval rather than from reasoning over correctly retrieved text.
Get each breakdown before it makes the rounds
You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.
No spam. Unsubscribe anytime.