[论文解读] Underneath the Numbers: Quantitative and Qualitative Gender Fairness in LLMs for Depression Prediction
本研究通过定量与定性方法,调查了大型语言模型(LLMs)在抑郁预测任务中的性别公平性。研究发现,尽管LLaMA 2在定量公平性指标上表现更优,但ChatGPT在生成解释时展现出更全面、连贯且与上下文相符的特性,凸显了LLMs在数值公平性与可解释性公平性之间的权衡。
Recent studies show bias in many machine learning models for depression detection, but bias in LLMs for this task remains unexplored. This work presents the first attempt to investigate the degree of gender bias present in existing LLMs (ChatGPT, LLaMA 2, and Bard) using both quantitative and qualitative approaches. From our quantitative evaluation, we found that ChatGPT performs the best across various performance metrics and LLaMA 2 outperforms other LLMs in terms of group fairness metrics. As qualitative fairness evaluation remains an open research question we propose several strategies (e.g., word count, thematic analysis) to investigate whether and how a qualitative evaluation can provide valuable insights for bias analysis beyond what is possible with quantitative evaluation. We found that ChatGPT consistently provides a more comprehensive, well-reasoned explanation for its prediction compared to LLaMA 2. We have also identified several themes adopted by LLMs to qualitatively evaluate gender fairness. We hope our results can be used as a stepping stone towards future attempts at improving qualitative evaluation of fairness for LLMs especially for high-stakes tasks such as depression detection.
研究动机与目标
- 调查现有LLMs——ChatGPT、LLaMA 2和Bard——在抑郁检测任务中的性别偏见。
- 探讨纯定量公平性评估在LLMs中的局限性。
- 开发并应用超越数值指标的定性公平性评估策略。
- 研究LLMs如何在心理健康预测任务中理解并回应公平性标准。
- 通过解释质量与连贯性,贡献一种以人为本的公平性评估框架。
提出的方法
- 在两个广泛使用的抑郁检测数据集上,对三种LLMs——ChatGPT、LLaMA 2和Bard——进行了对比评估。
- 应用了群体公平性、准确率和精确率等定量公平性指标,评估不同性别子群体的表现。
- 开发了定性评估策略,包括主题分析和词数分析,以评估解释质量。
- 基于解释的全面性、连贯性、具体性与一致性,评估LLMs在公平性相关推理中的表现。
- 采用以人为本的标准,如可理解性、个性化与上下文相关性,评估解释质量。
- 识别出LLM生成的公平性评估中反复出现的主题,如使用性别中立语言及避免性别刻板印象。

实验结果
研究问题
- RQ1现有LLMs在抑郁预测任务中在多大程度上表现出性别偏见?
- RQ2LLMs在不同性别子群体中的定量公平性表现有何差异?
- RQ3对解释的定性评估能否揭示超越数值指标的公平性洞见?
- RQ4LLMs在评估自身在心理健康预测中的性别公平性时,使用了哪些主题?
- RQ5解释质量与连贯性如何影响对LLM生成预测中公平性的感知?
主要发现
- LLaMA 2在定量公平性指标上优于其他LLMs,在不同性别子群体中的预测表现更为均衡。
- 与LLaMA 2相比,ChatGPT提供了更全面、连贯且与上下文相关的解释,而LLaMA 2的回应常出现不一致或自相矛盾的情况。
- LLaMA 2频繁完成用户请求或添加无关内容,而非回应公平性评估提示,表明其任务遵循能力较差。
- ChatGPT始终如一地使用性别中立代词如‘they’或‘participant’,符合公平性最佳实践。
- LLaMA 2批评使用‘participant’为非性别化表达,却与其自身声明的公平性标准相矛盾,暴露出内部不一致。
- 本研究揭示了一种权衡:LLaMA 2在定量公平性方面表现优异,而ChatGPT则通过解释质量展现出更优的定性公平性。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。