Skip to main content
QUICK REVIEW

[Paper Review] Unveiling and Mitigating Bias in Mental Health Analysis with Large Language Models

Yuqing Wang, Yun Zhao|arXiv (Cornell University)|Jun 17, 2024
Mental Health via WritingPsychology3 citations
TL;DR

This study systematically evaluates and mitigates bias in large language models (LLMs) for mental health analysis across seven social factors using ten LLMs and eight diverse datasets. It finds that GPT-4 achieves the best fairness-performance balance, larger models show reduced bias, and fairness-aware prompting effectively reduces disparities across demographic groups, especially for underrepresented identities like religion and nationality.

ABSTRACT

The advancement of large language models (LLMs) has demonstrated strong capabilities across various applications, including mental health analysis. However, existing studies have focused on predictive performance, leaving the critical issue of fairness underexplored, posing significant risks to vulnerable populations. Despite acknowledging potential biases, previous works have lacked thorough investigations into these biases and their impacts. To address this gap, we systematically evaluate biases across seven social factors (e.g., gender, age, religion) using ten LLMs with different prompting methods on eight diverse mental health datasets. Our results show that GPT-4 achieves the best overall balance in performance and fairness among LLMs, although it still lags behind domain-specific models like MentalRoBERTa in some cases. Additionally, our tailored fairness-aware prompts can effectively mitigate bias in mental health predictions, highlighting the great potential for fair analysis in this field.

Motivation & Objective

  • To investigate fairness disparities in LLMs across diverse social factors such as gender, age, religion, and sexuality in mental health prediction.
  • To evaluate the performance and fairness of ten LLMs—including general-purpose and domain-specific models—across eight mental health datasets.
  • To develop and validate fairness-aware prompting strategies that mitigate bias in LLM-based mental health analysis.
  • To examine the relationship between model scale and fairness, challenging the conventional performance-fairness trade-off.
  • To provide actionable insights for ethical deployment of LLMs in high-stakes mental health applications.

Proposed method

  • Demographic enrichment: Injected social factors (e.g., gender, religion) into LLM prompts to create 60 unique variations per data sample for bias evaluation.
  • Evaluated ten LLMs (e.g., GPT-4, Llama3, MentaLLama) across eight mental health datasets using zero-shot standard prompting and few-shot Chain-of-Thought (CoT) prompting.
  • Proposed fairness-aware prompts based on observed bias patterns to actively mitigate disparities in model predictions.
  • Conducted aggregated and stratified evaluations to measure performance and fairness across demographic subgroups.
  • Used clinical fairness metrics such as Equal Opportunity (EO) scores to assess model equity.
  • Performed manual error analysis to identify persistent issues like sentiment misjudgment and ambiguity in LLM outputs.

Experimental results

Research questions

  • RQ1To what extent do LLMs exhibit biased predictions across diverse demographic groups in mental health analysis?
  • RQ2How does model size influence fairness and performance in mental health prediction tasks?
  • RQ3Can few-shot Chain-of-Thought prompting improve both performance and fairness in LLM-based mental health analysis?
  • RQ4How effective are fairness-aware prompts in reducing bias across different LLMs and demographic subgroups?
  • RQ5What are the persistent limitations of LLMs in handling sensitive mental health content, especially for underrepresented identities?

Key findings

  • GPT-4 achieved the best overall balance between performance and fairness among general-purpose LLMs, though it still underperforms domain-specific models like MentalRoBERTa on certain tasks.
  • Larger LLMs demonstrated lower bias levels, challenging the traditional performance-fairness trade-off and suggesting scale enhances fairness through better representation of diverse groups.
  • Few-shot Chain-of-Thought prompting improved both performance and fairness, indicating that reasoning and contextual cues enhance model robustness in mental health contexts.
  • Fairness-aware prompts significantly reduced bias across all evaluated LLMs, regardless of size, demonstrating the effectiveness of targeted prompting strategies for bias mitigation.
  • LLMs showed stronger performance for gender and age but struggled with religion and nationality, indicating persistent disparities for less represented identities.
  • Manual error analysis revealed recurring issues such as sentiment misjudgment and ambiguity, especially in texts with complex emotional or cultural nuances.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.