Skip to main content
QUICK REVIEW

[论文解读] Language Shapes Mental Health Evaluations in Large Language Models

Jiayi Xu, Xiyang Hu|arXiv (Cornell University)|Mar 6, 2026
Mental Health Treatment and Access被引用 0
一句话总结

论文表明对于 GPT-4o 和 Qwen3,用中文提示比英文提示在污名相关的评估取向上更高,并在后续精神健康分类上产生偏移,显示语言相关偏见在大语言模型评估中的存在。

ABSTRACT

This study investigates whether large language models (LLMs) exhibit cross-linguistic differences in mental health evaluations. Focusing on Chinese and English, we examine two widely used models, GPT-4o and Qwen3, to assess whether prompt language systematically shifts mental health-related evaluations and downstream decision outcomes. First, we assess models' evaluative orientation toward mental health stigma using multiple validated measurement scales capturing social stigma, self-stigma, and professional stigma. Across all measures, both models produce higher stigma-related responses when prompted in Chinese than in English. Second, we examine whether these differences also manifest in two common downstream decision tasks in mental health. In a binary mental health stigma detection task, sensitivity to stigmatizing content varies across language prompts, with lower sensitivity observed under Chinese prompts. In a depression severity classification task, predicted severity also differs by prompt language, with Chinese prompts associated with more underestimation errors, indicating a systematic downward shift in predicted severity relative to English prompts. Together, these findings suggest that language context can systematically shape evaluative patterns in LLM outputs and shift decision thresholds in downstream tasks.

研究动机与目标

  • 评估跨语言差异(中文 vs. 英文)是否系统性地影响大语言模型输出中的精神健康污名评估。
  • 确定语言提示是否影响下游决策任务,如污名检测和抑郁严重程度分类。
  • 将构建层面的污名模式与下游决策阈值相关联,以评估在不同语言下的潜在公平性影响。

提出的方法

  • 通过官方 API 对两个多语言大模型(GPT-4o 和 Qwen3-32B)进行查询,温度设为 0.0 以获得确定性输出。
  • 使用经过验证的心理测量学污名量表,覆盖社会、自我和专业领域,评估评估取向。
  • 包括小故事情境和基于 DSM 的情境,以在污名评估中捕捉社交距离和感知危险性。
  • 翻译并对齐一份中文字版污名检测小故事,使在零样本设置下实现成对跨语言比较。
  • 使用成对语言提示评估两个下游任务(二元污名检测和四级抑郁严重程度分类),每个样本聚合 30 次运行。

实验结果

研究问题

  • RQ1GPT-4o 和 Qwen3 在用中文 vs. 英文提示时,是否在精神健康污名评估取向上存在跨语言差异?
  • RQ2语言诱发的评估差异是否会转化为下游污名检测和抑郁严重程度分类的可度量差异?
  • RQ3提示语言在不同模型与任务中对预测的校准偏移的方向与幅度是多少?
  • RQ4观察到的跨语言效应是否对抽样变异性(温度)具有鲁棒性,并且在多种污名量表中是否一致?

主要发现

  • 两种模型在社会、自我和专业领域的污名相关分数在中文提示下普遍更高。
  • GPT-4o 在中文提示下表现出更高的感知公共污名(DDS)和个人污名(MISS),Qwen3 也呈现类似模式(补充数据)。
  • 中文提示提升了 GPT-4o 和 Qwen3 的抑郁相关污名(感知与个人)。
  • 在自我污名(SSOSH)和健康专业人员污名(OMS-HC)方面,中文提示下两模型均更高。
  • 在污名检测任务中,英文提示下 GPT-4o 的准确率更高(0.737 对 0.717),Qwen3 的准确率亦然(0.769 对 0.722),英文在召回率上的提升更明显,尤其是 Qwen3。
  • 在抑郁严重程度分类中,总体准确率差异较小,但中文提示下存在更多低估(例如 GPT-4o 为 43 对 11 的低估;Qwen3 为 46 对 17)。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。