[论文解读] Better Call GPT, Comparing Large Language Models Against Lawyers
高级大型语言模型与初级律师及LPOs进行比较,以资深律师作为基准,显示在合同审查方面具有可比的准确性、更快的审查时间以及成本显著降低。
This paper presents a groundbreaking comparison between Large Language Models and traditional legal contract reviewers, Junior Lawyers and Legal Process Outsourcers. We dissect whether LLMs can outperform humans in accuracy, speed, and cost efficiency during contract review. Our empirical analysis benchmarks LLMs against a ground truth set by Senior Lawyers, uncovering that advanced models match or exceed human accuracy in determining legal issues. In speed, LLMs complete reviews in mere seconds, eclipsing the hours required by their human counterparts. Cost wise, LLMs operate at a fraction of the price, offering a staggering 99.97 percent reduction in cost over traditional methods. These results are not just statistics, they signal a seismic shift in legal practice. LLMs stand poised to disrupt the legal industry, enhancing accessibility and efficiency of legal services. Our research asserts that the era of LLM dominance in legal contract review is upon us, challenging the status quo and calling for a reimagined future of legal workflows.
研究动机与目标
- 评估LLMs是否能够在合同中定位和判定法律问题方面优于初级律师和LPO。
- 评估LLMs相对于人类从业者在合同审查中的速度。
- 评估基于LLM的合同审查相较于人类从业者的成本。
- 将多种知名LLM在真实采购合同中以资深律师为基准进行基准测试。
提出的方法
- 以十份匿名采购合同(美国和新西兰)作为数据集。
- 通过资深律师的判定和问题位置确立基准答案。
- 使用精确率、召回率、F1分数和损失,将LLMs、初级律师和LPOs与基准答案进行比较。
- 分析各组每份文档的时间和成本。
- 选择具有大上下文窗口(>=16,000标记)的模型,并报告其设置和提示。

实验结果
研究问题
- RQ1LLMs在合同中对法律问题的判定和定位方面是否优于初级律师和LPOs?
- RQ2LLMs是否能比初级律师和LPOs更快地审查合同?
- RQ3LLMs在合同审查方面是否比初级律师和法律流程外包商更便宜?
主要发现
- LLMs(例如GPT4-1106)在问题判定的F分数约为0.87,接近LPOs,略高于初级律师。
- LLMs在问题定位的表现因模型而异,GPT4-32k的F-score约为0.74,而GPT4-1106达到0.69。
- 每份文档时间:Palm2 text-bison 0.73分钟;GPT-1106 4.7分钟;人类视角色而定,范围43–201分钟。
- 每份文档成本:LLMs约0.02美元到2.50美元/文档,远低于人工审阅者(如初级律师约74美元/文档,资深律师约76美元)。
- LLMs显示出显著的高效性和成本优势潜力,请在任务强调问题判定还是定位时慎重选择模型。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。