[Paper Review] LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models
LegalBench introduces a collaboratively constructed, open-source benchmark of 162 legal reasoning tasks across six reasoning types to evaluate LLMs, with an interdisciplinary construction process and initial empirical evaluation of 20 models.
The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. To enable cross-disciplinary conversations about LLMs in the law, we additionally show how popular legal frameworks for describing legal reasoning -- which distinguish between its many forms -- correspond to LegalBench tasks, thus giving lawyers and LLM developers a common vocabulary. This paper describes LegalBench, presents an empirical evaluation of 20 open-source and commercial LLMs, and illustrates the types of research explorations LegalBench enables.
Motivation & Objective
- Motivate the need for rigorous, domain-aligned benchmarks for legal reasoning in LLMs.
- Present a typology of legal reasoning grounded in IRAC and legal practice.
- Describe the construction, documentation, and collaborative process behind LegalBench.
- Provide an initial empirical evaluation of multiple LLMs across diverse task types and prompts.
- Offer a platform to enable further interdisciplinary research and practical application in legal AI.
Proposed method
- Introduce a six-type legal reasoning typology (issue-spotting, rule-recall, rule-application, rule-conclusion, interpretation, rhetorical-understanding).
- Assemble 162 tasks from 36 data sources, including hand-crafted datasets by legal professionals and restructured existing corpora.
- Organize tasks with documentation, base prompts, and evaluation protocols to enable replicability.
- Evaluate 20 LLMs from 11 families across sizes using standardized prompts and prompts engineering strategies.
- Provide an answer-guide and multi-faceted evaluation (correctness and analysis) for rule-application tasks.
- Discuss limitations, interoperability with IRAC, and implications for policy, safety, and future work.

Experimental results
Research questions
- RQ1What types of legal reasoning can LLMs perform and how can they be measured in a fine-grained, domain-aligned benchmark?
- RQ2How can a collaborative, domain-expert-driven process improve the relevance and utility of LLM evaluations in law?
- RQ3How do different LLMs perform across a detailed typology of legal tasks and prompt strategies?
- RQ4To what extent can LegalBench tasks be extended to non-American jurisdictions and longer documents?
Key findings
- LegalBench provides 162 tasks spanning six reasoning types drawn from legal frameworks and practice.
- The benchmark enables standardized prompts, demonstrations, and evaluation protocols to study LLM performance in legal contexts.
- Initial experiments across 20 LLMs show varying strengths across task types and reveal insights into prompt-engineering strategies (details in the paper).
- LegalBench highlights the importance of domain-expert input in task construction to ensure practically useful and interpretable evaluations.
- There is an intentional emphasis on interpretive and contract-related tasks due to their ubiquitous legal language and practical implications.
- The authors discuss limitations (e.g., focus on English and American law, short context windows) and outline directions for future expansion.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.