Skip to main content
QUICK REVIEW

[論文レビュー] Better Call GPT, Comparing Large Language Models Against Lawyers

Lauren Martin, Nick Whitehouse|arXiv (Cornell University)|Jan 24, 2024
Artificial Intelligence in Law被引用数 9
ひとこと要約

高度なLLMは、Senior LawyersをグラウンドトゥルースとしてJunior LawyersおよびLPOと比較され、契約レビューにおいて精度は同等、レビュー時間ははるかに短く、コストは劇的に低減されることを示しています。

ABSTRACT

This paper presents a groundbreaking comparison between Large Language Models and traditional legal contract reviewers, Junior Lawyers and Legal Process Outsourcers. We dissect whether LLMs can outperform humans in accuracy, speed, and cost efficiency during contract review. Our empirical analysis benchmarks LLMs against a ground truth set by Senior Lawyers, uncovering that advanced models match or exceed human accuracy in determining legal issues. In speed, LLMs complete reviews in mere seconds, eclipsing the hours required by their human counterparts. Cost wise, LLMs operate at a fraction of the price, offering a staggering 99.97 percent reduction in cost over traditional methods. These results are not just statistics, they signal a seismic shift in legal practice. LLMs stand poised to disrupt the legal industry, enhancing accessibility and efficiency of legal services. Our research asserts that the era of LLM dominance in legal contract review is upon us, challenging the status quo and calling for a reimagined future of legal workflows.

研究の動機と目的

  • LLMが契約の法的問題の所在と判定において、Junior LawyersおよびLPOより優れているかを評価する。
  • LLMと人間実務家の契約レビューの速度を評価する。
  • LLMベースの契約レビューのコストを人間実務家と比較して評価する。
  • 現実の調達契約においてSenior Lawyersをグランドトゥルースとして、複数の著名なLLMをベンチマークする。

提案手法

  • データセットとして10件の匿名化された調達契約(米国およびニュージーランド)を使用する。
  • Senior Lawyersの判断と問題の所在をグランドトゥルースとして確立する。
  • LLMs、Junior Lawyers、LPOをグランドトゥルースと比較し、精度・再現率・Fスコア・ロスを評価する。
  • 各グループの文書あたりの所要時間とコストを分析する。
  • 大きなコンテキスト窓(>=16,000トークン)を持つモデルを選択し、その設定とプロンプトを報告する。
Figure 1. Level of agreement on issues by role
Figure 1. Level of agreement on issues by role

実験結果

リサーチクエスチョン

  • RQ1LLMsは契約における法的問題の判定と所在において、Junior LawyersおよびLPOを上回るか。
  • RQ2LLMsはJunior LawyersおよびLPOより契約を速くレビューできるか。
  • RQ3LLMsはJunior LawyersおよびLegal Process Outsourcersより契約を安くレビューできるか。

主な発見

  • LLMs(例:GPT4-1106)は、問題判定のFスコアを約0.87付近で達成し、LPOに匹敵し、Junior Lawyersよりわずかに上回る。
  • 問題の所在のパフォーマンスはモデルによって異なり、GPT4-32kは約0.74のFスコアを達成し、GPT4-1106は0.69に達する。
  • 文書あたりの所要時間:Palm2 text-bison 0.73分;GPT-1106 4.7分;人間は役割に応じて43–201分の範囲。
  • 文書あたりのコスト:LLMsは約0.02~2.50ドルの範囲であり、人間のレビュアー(例:Junior Lawyers約74ドル/文書、Senior Lawyers約76ドル)よりはるかに安い。
  • LLMsには劇的な効率化とコスト優位の可能性がある一方、問題判定とローカライズのどちらを重視するかに応じてモデルの慎重な選択が必要である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。