Skip to main content
QUICK REVIEW

[Paper Review] Auditing the Use of Language Models to Guide Hiring Decisions

Johann D. Gaebler, Sharad Goel|arXiv (Cornell University)|Apr 3, 2024
Multi-Agent Systems and Negotiation4 citations
TL;DR

This paper proposes using correspondence experiments—commonly used to detect human bias—in auditing large language models (LLMs) for racial and gender bias in hiring. By systematically varying names and pronouns in real K-12 teaching job applications, the authors find moderate disparities favoring women and racial minorities across multiple LLMs, suggesting that demographic inference persists even when explicit identifiers are removed.

ABSTRACT

Regulatory efforts to protect against algorithmic bias have taken on increased urgency with rapid advances in large language models (LLMs), which are machine learning models that can achieve performance rivaling human experts on a wide array of tasks. A key theme of these initiatives is algorithmic "auditing," but current regulations -- as well as the scientific literature -- provide little guidance on how to conduct these assessments. Here we propose and investigate one approach for auditing algorithms: correspondence experiments, a widely applied tool for detecting bias in human judgements. In the employment context, correspondence experiments aim to measure the extent to which race and gender impact decisions by experimentally manipulating elements of submitted application materials that suggest an applicant's demographic traits, such as their listed name. We apply this method to audit candidate assessments produced by several state-of-the-art LLMs, using a novel corpus of applications to K-12 teaching positions in a large public school district. We find evidence of moderate race and gender disparities, a pattern largely robust to varying the types of application material input to the models, as well as the framing of the task to the LLMs. We conclude by discussing some important limitations of correspondence experiments for auditing algorithms.

Motivation & Objective

  • To investigate whether large language models (LLMs) produce racially and gendered disparities in hiring assessments.
  • To evaluate the effectiveness of correspondence experiments as a tool for auditing algorithmic bias in LLM-based HR systems.
  • To test the robustness of observed disparities across variations in input materials, model instructions, and anonymization techniques.
  • To examine whether LLMs can infer demographic traits from anonymized application materials, undermining fairness-through-ignorance approaches.
  • To inform regulatory and policy efforts by providing a practical, empirically grounded method for auditing LLMs in high-stakes employment decisions.

Proposed method

  • Constructed a novel corpus of real job applications to K-12 teaching positions in a large Texas public school district, including resumes and video interview responses.
  • Employed state-of-the-art open-source and proprietary LLMs to generate hiring recommendations based on application materials.
  • Applied correspondence experiments by systematically varying applicant names and pronouns to simulate different racial and gender identities.
  • Conducted robustness checks by altering model instructions, inputting only resumes, and replacing school district references with a predominantly White district.
  • Used a control condition where names and pronouns were removed from inputs to test the efficacy of 'fairness through unawareness'.
  • Analyzed model outputs for disparities in hiring recommendations across demographic groups, focusing on score differentials.

Experimental results

Research questions

  • RQ1Do state-of-the-art LLMs exhibit measurable race and gender disparities in hiring recommendations for K-12 teaching positions?
  • RQ2How robust are these disparities to variations in input materials (e.g., resumes only vs. full dossiers) and model instructions?
  • RQ3To what extent can LLMs infer demographic attributes from anonymized application materials, undermining fairness-through-ignorance approaches?
  • RQ4How do model outputs change when the school district context is altered to a predominantly White district?
  • RQ5Can correspondence experiments serve as a reliable method for auditing LLMs for bias in employment settings?

Key findings

  • The study found moderate race and gender disparities in LLM-generated hiring recommendations, with women and racial minorities receiving higher scores than White men.
  • These disparities were robust across different model instructions, input types (e.g., resumes only), and variations in the school district context.
  • Even when names and pronouns were removed from input materials, LLMs still produced disparate assessments, indicating that demographic inference occurs through implicit cues.
  • The model's ability to infer demographics from anonymized inputs suggests that fairness-through-ignorance is insufficient to prevent bias.
  • Disparities persisted when the school district was replaced with a predominantly White district in West Virginia, indicating that the effect is not driven by local demographic context.
  • The results demonstrate that correspondence experiments can detect algorithmic bias in LLMs, offering a practical tool for regulatory auditing.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.