[Paper Review] Language Shapes Mental Health Evaluations in Large Language Models
The paper shows that for GPT-4o and Qwen3, prompting in Chinese yields higher stigma-related evaluative orientations and shifts downstream mental health classifications compared to English, indicating language-contingent biases in LLM assessments.
This study investigates whether large language models (LLMs) exhibit cross-linguistic differences in mental health evaluations. Focusing on Chinese and English, we examine two widely used models, GPT-4o and Qwen3, to assess whether prompt language systematically shifts mental health-related evaluations and downstream decision outcomes. First, we assess models' evaluative orientation toward mental health stigma using multiple validated measurement scales capturing social stigma, self-stigma, and professional stigma. Across all measures, both models produce higher stigma-related responses when prompted in Chinese than in English. Second, we examine whether these differences also manifest in two common downstream decision tasks in mental health. In a binary mental health stigma detection task, sensitivity to stigmatizing content varies across language prompts, with lower sensitivity observed under Chinese prompts. In a depression severity classification task, predicted severity also differs by prompt language, with Chinese prompts associated with more underestimation errors, indicating a systematic downward shift in predicted severity relative to English prompts. Together, these findings suggest that language context can systematically shape evaluative patterns in LLM outputs and shift decision thresholds in downstream tasks.
Motivation & Objective
- Assess whether cross-linguistic differences (Chinese vs. English) systematically shape mental health stigma evaluations in LLM outputs.
- Determine if language prompts influence downstream decision tasks such as stigma detection and depression severity classification.
- Link construct-level stigma patterns to downstream decision thresholds to assess potential fairness implications across languages.
Proposed method
- Query two multilingual LLMs (GPT-4o and Qwen3-32B) via official APIs with temperature set to 0.0 for deterministic outputs.
- Use validated psychometric stigma measures across social, self, and professional domains to assess evaluative orientation.
- Include vignette- and DSM-based scenarios to capture social distance and perceived dangerousness in stigma assessments.
- Translate and align a Chinese version of the stigma detection vignette to enable paired cross-linguistic comparisons in zero-shot settings.
- Evaluate two downstream tasks (binary stigma detection and four-level depression severity classification) using paired-language prompts and aggregated 30 runs per sample.
Experimental results
Research questions
- RQ1Do GPT-4o and Qwen3 exhibit cross-linguistic differences in mental health stigma evaluative orientation when prompted in Chinese vs. English?
- RQ2Do language-induced evaluative differences translate into measurable differences in downstream stigma detection and depression severity classification?
- RQ3What is the direction and magnitude of any calibration shifts in predictions caused by prompt language across models and tasks?
- RQ4Are observed cross-linguistic effects robust to sampling variability (temperature) and consistent across multiple stigma measures?
Key findings
- Both models produce higher stigma-related scores when prompted in Chinese across social, self, and professional domains.
- GPT-4o shows higher perceived public stigma (DDS) and personal stigma (MISS) with Chinese prompts; Qwen3 shows similar patterns (supplementary data).
- Chinese prompts increase depression-specific stigma (perceived and personal) for GPT-4o and Qwen3.
- Self-stigma (SSOSH) and health professional stigma (OMS-HC) are higher under Chinese prompts for both models.
- In stigma detection, English prompts yield higher accuracy for GPT-4o (0.737 vs 0.717) and Qwen3 (0.769 vs 0.722), with larger recall gains for English, especially for Qwen3.
- In depression severity classification, overall accuracy differences are small, but Chinese prompts show more under-estimation (e.g., 43 vs 11 under-estimations for GPT-4o; 46 vs 17 for Qwen3).
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.