Skip to main content
QUICK REVIEW

[Paper Review] A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models

Stephen Pfohl, Heather Cole-Lewis|arXiv (Cornell University)|Mar 18, 2024
Healthcare Systems and Public Health6 citations
TL;DR

The paper presents a multifactorial human evaluation framework and seven EquityMedQA datasets to surface health equity harms and biases in LLMs, demonstrated through a large Med-PaLM 2 case study.

ABSTRACT

Large language models (LLMs) hold promise to serve complex health information needs but also have the potential to introduce harm and exacerbate health disparities. Reliably evaluating equity-related model failures is a critical step toward developing systems that promote health equity. We present resources and methodologies for surfacing biases with potential to precipitate equity-related harms in long-form, LLM-generated answers to medical questions and conduct a large-scale empirical case study with the Med-PaLM 2 LLM. Our contributions include a multifactorial framework for human assessment of LLM-generated answers for biases, and EquityMedQA, a collection of seven datasets enriched for adversarial queries. Both our human assessment framework and dataset design process are grounded in an iterative participatory approach and review of Med-PaLM 2 answers. Through our empirical study, we find that our approach surfaces biases that may be missed via narrower evaluation approaches. Our experience underscores the importance of using diverse assessment methodologies and involving raters of varying backgrounds and expertise. While our approach is not sufficient to holistically assess whether the deployment of an AI system promotes equitable health outcomes, we hope that it can be leveraged and built upon towards a shared goal of LLMs that promote accessible and equitable healthcare.

Motivation & Objective

  • Define dimensions of bias with potential to cause equity-related harms in medical LLM outputs via participatory, expert-driven design.
  • Develop three rubrics (independent, pairwise, counterfactual) for structured human evaluation of bias in long-form medical answers.
  • Create EquityMedQA: seven adversarial datasets to probe health equity biases in LLMs.
  • Demonstrate applicability through a large-scale empirical study on Med-PaLM and Med-PaLM 2 with diverse raters and rich qualitative insights.

Proposed method

  • Iterative, participatory rubric design with Equity AI experts (EARR) and physicians to identify bias dimensions.
  • Design of three evaluation rubrics (independent, pairwise, counterfactual) aligned with six bias dimensions.
  • Release of EquityMedQA: seven datasets combining human-curated and LLM-generated adversarial queries (total 4,668 examples).
  • Large-scale empirical study applying rubrics across Med-PaLM and Med-PaLM 2 outputs, with 806 raters from clinicians, equity experts, and consumers.
  • Quantitative and qualitative analysis of >17,000 human ratings to assess inter-rater reliability and bias explanations.

Experimental results

Research questions

  • RQ1What bias dimensions in LLM-generated medical answers most contribute to equity-related harms?
  • RQ2Can a multifactorial, participatory rubric reveal biases missed by traditional evaluation approaches?
  • RQ3Do adversarial, equity-focused datasets uncover vulnerabilities in Med-PaLM 2 not seen with standard datasets?
  • RQ4How does rater diversity affect detection and interpretation of health equity biases in LLM outputs?

Key findings

  • A diverse rater pool (806 individuals) and multiple assessment rubrics reveal biases that narrower evaluations miss.
  • Combining seven EquityMedQA datasets with a multi-rubric evaluation uncovers equity-related harms not captured by single-dataset approaches.
  • The open, adversarial data and counterfactual designs help differentiate contextually meaningful bias from non-meaningful changes.
  • Adversarial testing and participatory rubric design improve detection of several bias dimensions, including inaccurate portrayals, lack of inclusivity, and omission of structural explanations.
  • The authors emphasize that bias identification alone does not certify equitable health outcomes, highlighting the need for broader, context-specific evaluations.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.