[论文解读] Gendered Mental Health Stigma in Masked Language Models
本文通過基於心理學的提示語,調查掩碼語言模型(MLMs)中的性別化心理健康污名,以評估模型在生成與心理健康狀況相關的性別詞彙時的偏見。研究發現,MLMs 對女性的心理健康問題存在不成比例的關聯——特別是在尋求治療的情境中——而對男性則存在逃避與缺乏幫助的刻板印象;然而,經過心理健康數據微調的模型在一定程度上減少了這些偏見。
Mental health stigma prevents many individuals from receiving the appropriate care, and social psychology studies have shown that mental health tends to be overlooked in men. In this work, we investigate gendered mental health stigma in masked language models. In doing so, we operationalize mental health stigma by developing a framework grounded in psychology research: we use clinical psychology literature to curate prompts, then evaluate the models' propensity to generate gendered words. We find that masked language models capture societal stigma about gender in mental health: models are consistently more likely to predict female subjects than male in sentences about having a mental health condition (32% vs. 19%), and this disparity is exacerbated for sentences that indicate treatment-seeking behavior. Furthermore, we find that different models capture dimensions of stigma differently for men and women, associating stereotypes like anger, blame, and pity more with women with mental health conditions than with men. In showing the complex nuances of models' gendered mental health stigma, we demonstrate that context and overlapping dimensions of identity are important considerations when assessing computational models' social biases.
研究动机与目标
- 調查掩碼語言模型如何編碼性別化心理健康污名,特別是社會傾向忽視男性心理健康的心理機制。
- 評估語言模型是否會放大或減輕將心理疾病與特定性別聯繫起來的社會刻板印象。
- 評估針對特定領域的預訓練(在心理健康語料上)對減少或強化此類偏見的影響。
- 開發基於心理學的框架,使用臨床驗證的提示語,探測自然語言處理模型中的心理健康污名。
- 強調在心理健康應用中負責部署模型時,語境與交叉身份的重要性。
提出的方法
- 基於健康行為過程模型(HAPA)設計一組提示語,代表心理健康狀況的三個階段:診斷、意圖與行動。
- 使用帶有掩碼標記的提示語,評估模型在生成與心理健康相關句子中主語的性別詞彙(代詞、姓名、名詞)時的表現。
- 對比通用 MLM 與在心理健康語料上微調的模型,以評估訓練數據的影響。
- 基於歸因問卷(AQ-27)設計提示語,探測與心理疾病相關的刻板印象特徵(如憤怒、同情、責備)。
- 應用遞歸啟發式方法,比較不同模型輸出中性別詞彙的綜合機率。
- 分析模型在 11 种常見心理健康診斷中的行為,因註釋限制,專注於二元性別關聯。
实验结果
研究问题
- RQ1RQ1:掩碼語言模型是否對某一生理性別的心理健康狀況有更強烈的關聯,特別是在不同求助階段(診斷、意圖、行動)中?
- RQ2RQ2:掩碼語言模型對患有心理健康問題的男性與女性,其與刻板印象特徵(如責備、同情、憤怒)的關聯有何差異?
- RQ3RQ3:在心理健康語料上預訓練是否會影響模型的性別化心理健康污名,相較於通用模型?
- RQ4RQ4:模型生成的性別關聯在不同心理健康診斷與語境提示下是否具有一致性?
- RQ5RQ5:多重身份與語境的重疊如何影響語言模型中性別化心理健康污名的表現?
主要发现
- 與男性相比,掩碼語言模型在生成與心理健康狀況相關的句子時,有 32% 的機率生成與女性相關的詞彙,而男性僅為 19%。
- 當提示語顯示治療求助行為時,這種性別差異顯著擴大,行動階段提示中女性關聯上升至 38%。
- 在心理健康語料上微調的模型減少了主語生成中的性別差異,促進了更具性別中立的回應。
- 通用模型將負面刻板印象(如責備、同情、憤怒)更強烈地與患有心理健康問題的女性聯繫起來,而非男性。
- 相反,模型更強烈地將男性與逃避及缺乏求助行為聯繫起來,強化了男性情感沉默的污名。
- 研究顯示,不同模型在減少某些形式污名的同時,可能增加其他形式,突顯 NLP 系統中偏見的複雜性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。