[논문 리뷰] Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence
본 논문은 대형 언어 모델이 세법 관련 법적 추론을 어떻게 습득하고 개선하는지 조사하고, 모델 릴리스 전반에 걸친 출현하는 능력과 성능에 대한 맥락 및 프롬프트의 영향을 보여준다.
Better understanding of Large Language Models' (LLMs) legal analysis abilities can contribute to improving the efficiency of legal services, governing artificial intelligence, and leveraging LLMs to identify inconsistencies in law. This paper explores LLM capabilities in applying tax law. We choose this area of law because it has a structure that allows us to set up automated validation pipelines across thousands of examples, requires logical reasoning and maths skills, and enables us to test LLM capabilities in a manner relevant to real-world economic lives of citizens and companies. Our experiments demonstrate emerging legal understanding capabilities, with improved performance in each subsequent OpenAI model release. We experiment with retrieving and utilising the relevant legal authority to assess the impact of providing additional legal context to LLMs. Few-shot prompting, presenting examples of question-answer pairs, is also found to significantly enhance the performance of the most advanced model, GPT-4. The findings indicate that LLMs, particularly when combined with prompting enhancements and the correct legal texts, can perform at high levels of accuracy but not yet at expert tax lawyer levels. As LLMs continue to advance, their ability to reason about law autonomously could have significant implications for the legal profession and AI governance.
연구 동기 및 목표
- LLMs가 세법 분석을 어떻게 수행하는지 이해하고, 모델 진행에 따른 출현하는 능력을 식별한다.
- 법적 권한 및 맥락 정보를 제공하는 것이 LLM 성능에 미치는 영향을 평가한다.
- 가장 능력이 좋은 모델에서 few-shot 프롬프트의 효과를 평가한다.
- 현 시점의 LLM이 전문가 수준의 세법 추론에 도달할 수 있는지 판단한다.
제안 방법
- 수천 개의 세법 예제에 걸친 자동화된 검증 파이프라인을 구축하여 LLM 추론을 테스트한다.
- 관련 법적 권한을 검색하고 모델에 적절한 법적 맥락을 제공하기 위해 이를 통합한다.
- 출현하는 능력을 식별하기 위해 OpenAI 모델 릴리스(예: 이전 버전과 최신 버전) 간의 모델 성능을 비교한다.
- 질문-답변 쌍을 모델에 제시하여 few-shot 프롬프트를 평가한다.
- 추가 법령 텍스트와 맥락이 세법 문제 해결에 미치는 영향을 분석한다.
실험 결과
연구 질문
- RQ1모델이 릴리스가 진행될수록 세법에서 출현하는 법적 이해 능력을 보여주는가?
- RQ2법적 권한 및 맥락 제공이 세법 작업에서의 LLM 성능에 어떤 영향을 미치는가?
- RQ3가장 진보된 모델의 세법 추론 정확도를 few-shot 프롬프트가 크게 향상시키는가?
- RQ4현 시점의 LLM이 전문가급 세무 변호사 수준의 정확도와 일관성에 도달할 수 있는가?
주요 결과
- LLMs는 연속적인 모델 릴리스와 함께 법적 이해가 향상되는 모습을 보인다.
- 법적 권한 및 맥락 정보를 제공하면 성능이 향상된다.
- Few-shot 프롬프트는 가장 능력 있는 모델의 결과를 현저히 향상시킨다.
- LLMs는 세법 작업에서 높은 정확도를 달성할 수 있지만 아직 전문가급 세무 변호사 성능에는 도달하지 못한다.
- 향상된 LLM 기능은 법조계 및 AI 거버넌스에 의미 있는 시사점을 가질 수 있다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.