Skip to main content
QUICK REVIEW

[Paper Review] Large Language Models as Tax Attorneys: A Case Study in Legal Capabilities Emergence

John J. Nay, David Karamardian|arXiv (Cornell University)|Jun 12, 2023
Artificial Intelligence in LawSocial Sciences13 citations
TL;DR

The paper investigates how large language models acquire and improve legal reasoning in tax law, showing emergent capabilities across model releases and the impact of context and prompting on performance.

ABSTRACT

Better understanding of Large Language Models' (LLMs) legal analysis abilities can contribute to improving the efficiency of legal services, governing artificial intelligence, and leveraging LLMs to identify inconsistencies in law. This paper explores LLM capabilities in applying tax law. We choose this area of law because it has a structure that allows us to set up automated validation pipelines across thousands of examples, requires logical reasoning and maths skills, and enables us to test LLM capabilities in a manner relevant to real-world economic lives of citizens and companies. Our experiments demonstrate emerging legal understanding capabilities, with improved performance in each subsequent OpenAI model release. We experiment with retrieving and utilising the relevant legal authority to assess the impact of providing additional legal context to LLMs. Few-shot prompting, presenting examples of question-answer pairs, is also found to significantly enhance the performance of the most advanced model, GPT-4. The findings indicate that LLMs, particularly when combined with prompting enhancements and the correct legal texts, can perform at high levels of accuracy but not yet at expert tax lawyer levels. As LLMs continue to advance, their ability to reason about law autonomously could have significant implications for the legal profession and AI governance.

Motivation & Objective

  • Understand how LLMs perform tax law analysis and identify emergent capabilities across model progressions.
  • Evaluate the impact of providing legal authority and contextual information on LLM performance.
  • Assess the effect of few-shot prompting on the most capable model.
  • Determine whether current LLMs can reach expert-level tax-law reasoning.

Proposed method

  • Set up automated validation pipelines across thousands of tax-law examples to test LLM reasoning.
  • Retrieve and incorporate relevant legal authorities to provide proper legal context to the model.
  • Compare model performance across OpenAI model releases (e.g., earlier vs. latest) to identify emergent capabilities.
  • Evaluate few-shot prompting by presenting question-answer pairs to the models.
  • Analyze the effect of additional legal texts and context on tax-law problem solving.

Experimental results

Research questions

  • RQ1Do LLMs display emergent legal understanding capabilities in tax law as models advance over releases?
  • RQ2How does supplying legal authorities and context affect LLM performance on tax-law tasks?
  • RQ3Does few-shot prompting significantly boost the accuracy of the most advanced models in tax-law reasoning?
  • RQ4Can current LLMs reach expert tax-law attorney levels of accuracy and consistency?

Key findings

  • LLMs show improving legal understanding with successive model releases.
  • Providing legal authorities and contextual information enhances performance.
  • Few-shot prompting markedly improves the most capable model’s results.
  • LLMs can achieve high accuracy in tax-law tasks but do not yet reach expert-level tax attorney performance.
  • Advancing LLM capabilities could have meaningful implications for the legal profession and AI governance.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.