[论文解读] She Elicits Requirements and He Tests: Software Engineering Gender Bias in Large Language Models
本研究通过芬兰语进行回译,检测大型语言模型在56项软件工程任务中与隐性性别偏见的关系。结果显示,测试任务几乎完全与'他'关联(100%),而需求获取任务仅在6%的情况下与'他'关联,揭示了任务角色中显著的性别化关联,反映出并强化了社会刻板印象。
Implicit gender bias in software development is a well-documented issue, such as the association of technical roles with men. To address this bias, it is important to understand it in more detail. This study uses data mining techniques to investigate the extent to which 56 tasks related to software development, such as assigning GitHub issues and testing, are affected by implicit gender bias embedded in large language models. We systematically translated each task from English into a genderless language and back, and investigated the pronouns associated with each task. Based on translating each task 100 times in different permutations, we identify a significant disparity in the gendered pronoun associations with different tasks. Specifically, requirements elicitation was associated with the pronoun "he" in only 6% of cases, while testing was associated with "he" in 100% of cases. Additionally, tasks related to helping others had a 91% association with "he" while the same association for tasks related to asking coworkers was only 52%. These findings reveal a clear pattern of gender bias related to software development tasks and have important implications for addressing this issue both in the training of large language models and in broader society.
研究动机与目标
- 调查大型语言模型中的隐性性别偏见如何影响其与特定软件工程任务的关联。
- 通过数据挖掘技术识别哪些开发任务被不成比例地与男性代词'他'关联。
- 理解大型语言模型中性别化语言模式如何在软件工程角色中强化刻板印象。
- 为模型训练和任务分配中的针对性干预提供证据。
- 展示使用无性别语言进行回译作为检测自然语言处理系统中细微隐性性别偏见的方法的有效性。
提出的方法
- 回译:将每个软件开发任务首先从英语翻译成芬兰语(一种无性别语言),然后使用DeepL翻译器回译成英语。
- 该过程重复100次,任务列表采用随机排列,以最小化翻译中的上下文依赖偏见。
- 分析最终英语句子中的代词关联,以确定'他'、'她'、'他或她'、'他/她'及其他形式的出现频率。
- 本研究使用DeepL API进行翻译,利用其高质量输出检测模型生成文本中的隐性性别关联。
- 该方法基于假设:如果某项任务在回译后始终与'他'关联,则反映了模型训练数据中潜在的性别偏见。
- 将100次运行的结果汇总,以确保统计稳健性并减少单次翻译差异带来的噪声。
实验结果
研究问题
- RQ1大型语言模型中,不同软件工程任务与男性代词'他'的关联程度如何?
- RQ2在需求获取、测试和指导等任务中,性别代词关联如何变化?
- RQ3涉及沟通或协作的任务是否更可能与'她'或'他'关联?
- RQ4回译方法是否能有效揭示原始文本中不明显的大型语言模型中的隐性性别偏见?
- RQ5任务背景(如客户互动或技术自主性)在塑造代词关联方面发挥什么作用?
主要发现
- 在100%的回译中,测试任务与代词'他'相关联,表明该角色存在强烈且一致的男性偏见。
- 在仅6%的情况下,需求获取任务与'他'相关联,表明与男性代词的关联较弱或呈中性。
- 涉及帮助他人的任务在91%的情况下与'他'相关联,表明在支持性或协作性角色中存在显著的男性偏见。
- 与同事沟通的任务仅在52%的情况下与'他'相关联,表明与男性代词的关联明显较弱。
- 在全部100次回译中,注释和学习任务均与'他'相关联,凸显了在认知或知识共享任务中极端的偏见。
- 代词'她'更常与面向客户或沟通密集型任务(如指导和会议)相关联,表明在人际软件工程活动中存在性别化角色分配。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。