Skip to main content
QUICK REVIEW

[论文解读] A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models

Stephen Pfohl, Heather Cole-Lewis|arXiv (Cornell University)|Mar 18, 2024
Healthcare Systems and Public Health被引用 6
一句话总结

本论文提出一个多因素的人类评估框架和七个 EquityMedQA 数据集,以揭示健康公平性伤害和偏见在 LLMs 中的表现,这一结论通过对 Med-PaLM 2 的大规模案例研究来证明。

ABSTRACT

Large language models (LLMs) hold promise to serve complex health information needs but also have the potential to introduce harm and exacerbate health disparities. Reliably evaluating equity-related model failures is a critical step toward developing systems that promote health equity. We present resources and methodologies for surfacing biases with potential to precipitate equity-related harms in long-form, LLM-generated answers to medical questions and conduct a large-scale empirical case study with the Med-PaLM 2 LLM. Our contributions include a multifactorial framework for human assessment of LLM-generated answers for biases, and EquityMedQA, a collection of seven datasets enriched for adversarial queries. Both our human assessment framework and dataset design process are grounded in an iterative participatory approach and review of Med-PaLM 2 answers. Through our empirical study, we find that our approach surfaces biases that may be missed via narrower evaluation approaches. Our experience underscores the importance of using diverse assessment methodologies and involving raters of varying backgrounds and expertise. While our approach is not sufficient to holistically assess whether the deployment of an AI system promotes equitable health outcomes, we hope that it can be leveraged and built upon towards a shared goal of LLMs that promote accessible and equitable healthcare.

研究动机与目标

  • Define dimensions of bias with potential to cause equity-related harms in medical LLM outputs via participatory, expert-driven design.
  • Develop three rubrics (independent, pairwise, counterfactual) for structured human evaluation of bias in long-form medical answers.
  • Create EquityMedQA: seven adversarial datasets to probe health equity biases in LLMs.
  • Demonstrate applicability through a large-scale empirical study on Med-PaLM and Med-PaLM 2 with diverse raters and rich qualitative insights.

提出的方法

  • Iterative, participatory rubric design with Equity AI experts (EARR) and physicians to identify bias dimensions.
  • Design of three evaluation rubrics (independent, pairwise, counterfactual) aligned with six bias dimensions.
  • Release of EquityMedQA: seven datasets combining human-curated and LLM-generated adversarial queries (total 4,668 examples).
  • Large-scale empirical study applying rubrics across Med-PaLM and Med-PaLM 2 outputs, with 806 raters from clinicians, equity experts, and consumers.
  • Quantitative and qualitative analysis of >17,000 human ratings to assess inter-rater reliability and bias explanations.]
  • research_questions: ["What bias dimensions in LLM-generated medical answers most contribute to equity-related harms?", "Can a multifactorial, participatory rubric reveal biases missed by traditional evaluation approaches?", "Do adversarial, equity-focused datasets uncover vulnerabilities in Med-PaLM 2 not seen with standard datasets?", "How does rater diversity affect detection and interpretation of health equity biases in LLM outputs?"]
  • key_findings:[
  • ]},
  • table_headers: []
  • table_rows: []}
  • title_translated_in_target_language_not_required?

实验结果

研究问题

  • RQ1What bias dimensions in LLM-generated medical answers most contribute to equity-related harms?
  • RQ2Can a multifactorial, participatory rubric reveal biases missed by traditional evaluation approaches?
  • RQ3Do adversarial, equity-focused datasets uncover vulnerabilities in Med-PaLM 2 not seen with standard datasets?
  • RQ4How does rater diversity affect detection and interpretation of health equity biases in LLM outputs?

主要发现

  • A diverse rater pool (806 individuals) and multiple assessment rubrics reveal biases that narrower evaluations miss.
  • Combining seven EquityMedQA datasets with a multi-rubric evaluation uncovers equity-related harms not captured by single-dataset approaches.
  • The open, adversarial data and counterfactual designs help differentiate contextually meaningful bias from non-meaningful changes.
  • Adversarial testing and participatory rubric design improve detection of several bias dimensions, including inaccurate portrayals, lack of inclusivity, and omission of structural explanations.
  • The authors emphasize that bias identification alone does not certify equitable health outcomes, highlighting the need for broader, context-specific evaluations.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。