[论文解读] Understanding and Countering Stereotypes: A Computational Approach to the Stereotype Content Model
本文提出了一套计算框架,通过语义嵌入将文本中的社会刻板印象映射到刻板印象内容模型(Stereotype Content Model)的温暖与能力维度,并利用标注词典和基于调查的心理学研究验证其准确性。该研究进一步提出一种新颖的方法,用于生成反刻板印象——即反刻板印象的替代形式,可减少偏见思维,为自动化工具应对文本中的有害语言提供了基础。
Stereotypical language expresses widely-held beliefs about different social categories. Many stereotypes are overtly negative, while others may appear positive on the surface, but still lead to negative consequences. In this work, we present a computational approach to interpreting stereotypes in text through the Stereotype Content Model (SCM), a comprehensive causal theory from social psychology. The SCM proposes that stereotypes can be understood along two primary dimensions: warmth and competence. We present a method for defining warmth and competence axes in semantic embedding space, and show that the four quadrants defined by this subspace accurately represent the warmth and competence concepts, according to annotated lexicons. We then apply our computational SCM model to textual stereotype data and show that it compares favourably with survey-based studies in the psychological literature. Furthermore, we explore various strategies to counter stereotypical beliefs with anti-stereotypes. It is known that countering stereotypes with anti-stereotypical examples is one of the most effective ways to reduce biased thinking, yet the problem of generating anti-stereotypes has not been previously studied. Thus, a better understanding of how to generate realistic and effective anti-stereotypes can contribute to addressing pressing societal concerns of stereotyping, prejudice, and discrimination.
研究动机与目标
- 开发一种计算方法,将文本中的刻板印象映射到刻板印象内容模型(SCM)的温暖–能力维度的语义嵌入空间中。
- 通过与与这些特质相关的词典进行对比,验证模型在表示温暖和能力方面的准确性。
- 将计算得出的刻板印象预测结果与心理学调查研究的发现进行比较,证明其与既有文献的一致性。
- 探索并分析人类生成的反刻板印象,作为迈向自动化生成建设性反叙事的第一步。
- 通过提出一种计算框架,识别并应对由负面刻板印象及“积极”刻板印象造成的社会危害。
提出的方法
- 使用已知代表这些特质的标注词典中的种子词,在预训练的词嵌入空间中定义温暖与能力轴。
- 通过向量运算和余弦相似度,将 StereoSet 数据集中刻板印象和反刻板印象的短语投射到二维的温暖–能力平面上。
- 通过在温暖与能力词语的黄金标准词典上评估性能,优化嵌入模型的选择。
- 基于其在温暖–能力子空间中的位置对刻板印象进行聚类与解释,识别出高/低温暖与能力对应的象限。
- 分析 StereoSet 数据集中人类生成的反刻板印象,以理解其相对于刻板印象的语义与关系特性。
- 提出一种“算法在环”方法,以建议反刻板印象,强调人工监督以防止引入新的偏见。
实验结果
研究问题
- RQ1刻板印象内容模型的温暖与能力维度是否能在语义嵌入空间中有效表示并计算?
- RQ2计算得出的 StereoSet 数据集中刻板印象的温暖–能力得分,与传统心理学调查研究的发现有多吻合?
- RQ3人类生成的反刻板印象在语言和语义特征上有哪些特点?它们与原始刻板印象有何关系?
- RQ4能否系统性地生成反刻板印象以对抗文本中的有害或偏见关联?此类生成存在哪些风险?
- RQ5计算模型如何支持构建反叙事,以削弱基于群体的刻板印象,同时避免强化新的偏见?
主要发现
- 计算模型能够以高精度将刻板印象语言映射到温暖–能力平面上,经由温暖与能力词语的黄金标准词典验证。
- 模型的刻板印象预测结果与基于调查的心理学研究高度一致,证实其在反映现实社会认知方面的有效性。
- StereoSet 数据集中人类生成的反刻板印象主要位于其对应刻板印象的相反象限,表明存在系统性地反转负面关联的努力。
- 反刻板印象并非刻板印象的简单语义反面,而是具有语境和语义上的细微差别,通常反映原始刻板印象中不存在的替代性积极特质。
- 该方法计算效率高,仅需单个 CPU,在加载预训练嵌入后分析可在一分钟内完成。
- 本研究强调,即使出于良好意图的反刻板印象,若未经仔细筛选,也可能无意中强化新的偏见,凸显了人工在环系统的重要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。