Skip to main content
QUICK REVIEW

[논문 리뷰] Better Call GPT, Comparing Large Language Models Against Lawyers

Lauren Martin, Nick Whitehouse|arXiv (Cornell University)|2024. 01. 24.
Artificial Intelligence in Law인용 수 9
한 줄 요약

고급 LLM은 Senior Lawyers를 기준으로 ground truth로 삼아 Junior Lawyers 및 LPOs와 비교되며, 정확도는 비견되며 검토 시간이 훨씬 빠르고 계약 검토 비용이 극적으로 낮아진다는 것을 보여준다.

ABSTRACT

This paper presents a groundbreaking comparison between Large Language Models and traditional legal contract reviewers, Junior Lawyers and Legal Process Outsourcers. We dissect whether LLMs can outperform humans in accuracy, speed, and cost efficiency during contract review. Our empirical analysis benchmarks LLMs against a ground truth set by Senior Lawyers, uncovering that advanced models match or exceed human accuracy in determining legal issues. In speed, LLMs complete reviews in mere seconds, eclipsing the hours required by their human counterparts. Cost wise, LLMs operate at a fraction of the price, offering a staggering 99.97 percent reduction in cost over traditional methods. These results are not just statistics, they signal a seismic shift in legal practice. LLMs stand poised to disrupt the legal industry, enhancing accessibility and efficiency of legal services. Our research asserts that the era of LLM dominance in legal contract review is upon us, challenging the status quo and calling for a reimagined future of legal workflows.

연구 동기 및 목표

  • LLMs가 계약에서 법적 이슈를 찾고 판단하는 데 있어 Junior Lawyers와 LPOs를 능가할 수 있는지 평가한다.
  • 계약 검토에서 LLM의 속도를 인간 실무자와 비교하여 평가한다.
  • 인간 실무자와 비교한 LLM 기반 계약 검토 비용을 평가한다.
  • 실전 조달 계약에서 Senior Lawyers를 ground truth로 삼아 여러 주요 LLM을 벤치마크한다.

제안 방법

  • 데이터 세트로 미국(US) 및 뉴질랜드(NZ) 의익 비공개 조달 계약 10건을 사용한다.
  • Senior Lawyers의 판단 및 분쟁 위치를 통해 ground truth를 설정한다.
  • 정밀도, 재현율, F-점수 및 손실을 사용하여 LLM, Junior Lawyers, LPO를 ground truth와 대조한다.
  • 그룹별 문서당 시간 및 문서당 비용을 분석한다.
  • 대용량 컨텍스트 윈도우(>=16,000 토큰)를 가진 모델을 선택하고 설정과 프롬프트를 보고한다.
Figure 1. Level of agreement on issues by role
Figure 1. Level of agreement on issues by role

실험 결과

연구 질문

  • RQ1LLMs가 계약에서 법적 이슈의 판단 및 위치 파악에서 Junior Lawyers와 LPO를 능가하는가?
  • RQ2LLMs가 계약을 더 빠르게 검토할 수 있는가?
  • RQ3LLMs가 계약을 Junior Lawyers 및 Legal Process Outsourcers보다 비용 측면에서 더 저렴하게 검토할 수 있는가?

주요 결과

  • LLMs(예: GPT4-1106)는 이슈 결정(F-score)이 약 0.87 정도로 높고, LPO와 비슷하며 주니어 변호사보다 약간 높은 편이다.
  • LLM의 이슈 위치 탐지 성능은 모델에 따라 다르며, GPT4-32k는 약 0.74 F-score를 달성하고, GPT4-1106은 0.69에 도달한다.
  • 문서당 시간: Palm2 text-bison 0.73분; GPT-1106 4.7분; 인간은 역할에 따라 43–201분 범위.
  • 문서당 비용: LLM은 약 $0.02에서 $2.50 사이로 다양하며 인간 검토자에 비해 대단히 저렴하다(예: 주니어 변호사 ~$74 per doc, 시니어 변호사 ~$76).
  • LLMs는 이슈 결정 여부나 위치 파악 여부에 따라 과업의 강조점에 맞춰 신중한 모델 선택이 필요하다는 점에서 극적인 효율성 및 비용 이점을 보일 가능성을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.