Skip to main content
QUICK REVIEW

[Paper Review] Large Language Models: An Applied Econometric Framework

Jens Ludwig, Sendhil Mullainathan|arXiv (Cornell University)|Dec 9, 2024
Topic Modeling4 citations
TL;DR

This paper develops an econometric framework to assess when large language models (LLMs) can be reliably used in economics research. It distinguishes between prediction and estimation tasks, showing that LLMs are valid for prediction only if there is no training data leakage, and for estimation only if LLM outputs are as accurate as gold-standard measurements—otherwise, bias persists even with high accuracy.

ABSTRACT

Large language models (LLMs) enable researchers to analyze text at unprecedented scale and minimal cost. Researchers can now revisit old questions and tackle novel ones with rich data. We provide an econometric framework for realizing this potential in two empirical uses. For prediction problems -- forecasting outcomes from text -- valid conclusions require ``no training leakage'' between the LLM's training data and the researcher's sample, which can be enforced through careful model choice and research design. For estimation problems -- automating the measurement of economic concepts for downstream analysis -- valid downstream inference requires combining LLM outputs with a small validation sample to deliver consistent and precise estimates. Absent a validation sample, researchers cannot assess possible errors in LLM outputs, and consequently seemingly innocuous choices (which model, which prompt) can produce dramatically different parameter estimates. When used appropriately, LLMs are powerful tools that can expand the frontier of empirical economics.

Motivation & Objective

  • To establish rigorous conditions under which LLM outputs can be used in economic research without introducing bias.
  • To address the lack of formal econometric contracts for LLMs, analogous to those for traditional methods like OLS.
  • To identify when LLMs can be used for prediction versus estimation, and what assumptions are required for valid inference.
  • To provide practical guidance for researchers on avoiding common pitfalls in LLM-based empirical work.
  • To emphasize that LLMs cannot fully replace human data collection in experimental settings, especially when measuring human behavior.

Proposed method

  • Treat LLMs as black boxes, abstracting from their internal architecture and training data.
  • Distinguish between two types of empirical tasks: prediction (e.g., forecasting outcomes from text) and estimation (e.g., measuring economic concepts from text or human responses).
  • For prediction, require no training data leakage between the LLM’s training corpus and the researcher’s sample to ensure validity.
  • For estimation, require that LLM outputs are as accurate as the gold-standard measurement they replace, to avoid bias.
  • Propose collecting validation data to model LLM measurement error and correct for bias in estimation tasks.
  • Advocate for using open-source LLMs with documented training data and published weights to prevent leakage.

Experimental results

Research questions

  • RQ1Under what conditions can LLM outputs be used reliably for prediction in economic research?
  • RQ2When is it valid to use LLMs to estimate economic concepts derived from text or human responses?
  • RQ3How does training data leakage affect the validity of LLM-based predictions and estimates?
  • RQ4What role does measurement error in LLM outputs play in biasing econometric estimates?
  • RQ5In what research contexts can LLMs meaningfully simulate human subject responses without replacing real data collection?

Key findings

  • LLM use in prediction is valid only if there is no leakage between the LLM’s training data and the researcher’s sample.
  • For estimation tasks, LLM outputs must be as accurate as the gold-standard measurement to avoid bias, even if they are highly accurate.
  • Training data leakage is a major risk, especially when published experiments or surveys are part of the LLM’s training corpus.
  • LLM responses on economic reasoning tasks can be sensitive to prompt engineering, indicating brittleness in performance.
  • Validation data must be collected to model and correct for LLM measurement error in estimation tasks.
  • LLMs can only amplify, not fully replace, human subject data in experimental studies, especially when the number of experimental designs is large.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.