[Paper Review] Better Call GPT, Comparing Large Language Models Against Lawyers
Advanced LLMs are compared to Junior Lawyers and LPOs using Senior Lawyers as ground truth, showing comparable accuracy, much faster review times, and dramatically lower costs in contract review.
This paper presents a groundbreaking comparison between Large Language Models and traditional legal contract reviewers, Junior Lawyers and Legal Process Outsourcers. We dissect whether LLMs can outperform humans in accuracy, speed, and cost efficiency during contract review. Our empirical analysis benchmarks LLMs against a ground truth set by Senior Lawyers, uncovering that advanced models match or exceed human accuracy in determining legal issues. In speed, LLMs complete reviews in mere seconds, eclipsing the hours required by their human counterparts. Cost wise, LLMs operate at a fraction of the price, offering a staggering 99.97 percent reduction in cost over traditional methods. These results are not just statistics, they signal a seismic shift in legal practice. LLMs stand poised to disrupt the legal industry, enhancing accessibility and efficiency of legal services. Our research asserts that the era of LLM dominance in legal contract review is upon us, challenging the status quo and calling for a reimagined future of legal workflows.
Motivation & Objective
- Assess whether LLMs can outperform Junior Lawyers and LPOs in locating and determining legal issues in contracts.
- Evaluate the speed of LLMs versus human practitioners in contract review.
- Evaluate the cost of LLM-based contract review compared with human practitioners.
- Benchmark multiple prominent LLMs against Senior Lawyers as ground truth in real-world procurement contracts.
Proposed method
- Use ten anonymised procurement contracts (US and NZ) as the dataset.
- Establish ground truth via Senior Lawyers’ determinations and issue locations.
- Compare LLMs, Junior Lawyers, and LPOs against ground truth using precision, recall, F-score, and loss.
- Analyze time per document and cost per document across groups.
- Select models with large context windows (>=16,000 tokens) and report their settings and prompts.

Experimental results
Research questions
- RQ1Do LLMs outperform Junior Lawyers and LPOs in determination and location of legal issues in contracts?
- RQ2Can LLMs review contracts faster than Junior Lawyers and LPOs?
- RQ3Can LLMs review contracts cheaper than Junior Lawyers and Legal Process Outsourcers?
Key findings
- LLMs (e.g., GPT4-1106) achieve high issue-determination F-scores around 0.87, comparable to LPOs and slightly above Junior Lawyers.
- Issue-location performance by LLMs varies by model, with GPT4-32k achieving approximately 0.74 F-score, while GPT4-1106 reaches 0.69.
- Time per document: Palm2 text-bison 0.73 minutes; GPT-1106 4.7 minutes; humans range 43–201 minutes depending on role.
- Cost per document: LLMs range from about $0.02 to $2.50 per document, vastly cheaper than human reviewers (e.g., Junior Lawyers ~$74 per doc, Senior Lawyers ~$76).
- LLMs show potential for dramatic efficiency and cost advantages, with a need for careful model selection depending on whether the task emphasizes issue determination or localization.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.