Skip to main content
QUICK REVIEW

[论文解读] Gender bias in (non)-contextual clinical word embeddings for stereotypical medical categories

Gizem Soğancıoğlu, Fabian Mijsters|arXiv (Cornell University)|Aug 2, 2022
Sex and Gender in Healthcare被引用 5
一句话总结

本研究使用BioWordVec(非上下文相关)和clinical-BERT(上下文相关)在三种医学类别(精神障碍、性传播疾病和人格特质)中评估临床词嵌入中的性别偏见。研究发现,两种模型均表现出性别偏见,但BioWordVec的偏见显著更高,尤其在精神障碍方面;部分偏见与医学文献相悖,例如将抑郁症与男性关联,尽管女性的诊断率更高。

ABSTRACT

Clinical word embeddings are extensively used in various Bio-NLP problems as a state-of-the-art feature vector representation. Although they are quite successful at the semantic representation of words, due to the dataset - which potentially carries statistical and societal bias - on which they are trained, they might exhibit gender stereotypes. This study analyses gender bias of clinical embeddings on three medical categories: mental disorders, sexually transmitted diseases, and personality traits. To this extent, we analyze two different pre-trained embeddings namely (contextualized) clinical-BERT and (non-contextualized) BioWordVec. We show that both embeddings are biased towards sensitive gender groups but BioWordVec exhibits a higher bias than clinical-BERT for all three categories. Moreover, our analyses show that clinical embeddings carry a high degree of bias for some medical terms and diseases which is conflicting with medical literature. Having such an ill-founded relationship might cause harm in downstream applications that use clinical embeddings.

研究动机与目标

  • 调查在典型医学类别中预训练的临床词嵌入的性别偏见。
  • 比较上下文相关(clinical-BERT)与非上下文相关(BioWordVec)词嵌入中性别偏见的程度。
  • 确定嵌入中的偏见是否与既定医学文献中关于性别患病率的结论一致或冲突。
  • 分析MIMIC-III数据集中的人口统计不平衡,作为观察到的偏见的潜在根源。
  • 强调在下游临床机器学习应用中部署有偏见嵌入可能带来的伦理风险。

提出的方法

  • 在MIMIC-III临床笔记数据集上训练BioWordVec和clinical-BERT。
  • 使用词嵌入与性别相关术语(如“man”与“woman”)的相似性计算直接性别偏见得分。
  • 在三种医学类别(精神障碍、性传播疾病和人格特质)中分析偏见。
  • 通过MIMIC-III数据集的描述性分析,检查诊断中的性别分布。
  • 将偏见分类为“准确”(与医学文献一致)或“冲突”(与既定患病率数据矛盾)。
  • 应用标准化的偏见评分框架,比较BioWordVec与clinical-BERT之间的偏见水平。

实验结果

研究问题

  • RQ1上下文相关(clinical-BERT)与非上下文相关(BioWordVec)临床词嵌入之间的性别偏见水平如何比较?
  • RQ2临床词嵌入在多大程度上反映了或与已知的精神障碍、性传播疾病和人格特质的性别患病率模式一致?
  • RQ3在临床嵌入中,哪些医学术语表现出最高的性别偏见水平?
  • RQ4MIMIC-III数据集的人口统计分布如何导致词嵌入中的偏见?
  • RQ5对于下游临床NLP应用,冲突性偏见(例如将抑郁症与男性关联)有何影响?

主要发现

  • 在所有三个医学类别中,BioWordVec的平均直接性别偏见得分(0.06)高于clinical-BERT(0.03)。
  • 在精神障碍方面,BioWordVec的直接偏见得分为0.07,而clinical-BERT为0.04,表明某些疾病与男性有更强关联。
  • 双相情感障碍、精神分裂症和强迫症在两种嵌入中均对女性表现出偏见,尽管医学文献中并无显著性别差异。
  • 在两种模型中,抑郁症均表现出对男性的偏见,这与临床证据(女性诊断率更高)相矛盾。
  • MIMIC-III数据集显示,诊断为精神障碍的患者中男性比例更高,这导致了嵌入中的偏见。
  • 部分偏见(如‘anxiety’和‘breast cancer’)与医学文献一致,表明并非所有偏见都是错误或有害的。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。