[论文解读] What You See Is What You Get? The Impact of Representation Criteria on Human Bias in Hiring
本研究探讨了代表标准——候选人名单中固定比例的性别分布——如何影响招聘决策中的人类性别偏见。通过在亚马逊机械 Turk 上进行受控实验,研究发现,在世界分布严重失衡的职业中,平衡性别代表性可减轻偏见;但在人类偏好持续存在的职业中则无效,且决策者性别与任务复杂性进一步影响结果。
Although systematic biases in decision-making are widely documented, the ways in which they emerge from different sources is less understood. We present a controlled experimental platform to study gender bias in hiring by decoupling the effect of world distribution (the gender breakdown of candidates in a specific profession) from bias in human decision-making. We explore the effectiveness of \ extit{representation criteria}, fixed proportional display of candidates, as an intervention strategy for mitigation of gender bias by conducting experiments measuring human decision-makers' rankings for who they would recommend as potential hires. Experiments across professions with varying gender proportions show that balancing gender representation in candidate slates can correct biases for some professions where the world distribution is skewed, although doing so has no impact on other professions where human persistent preferences are at play. We show that the gender of the decision-maker, complexity of the decision-making task and over- and under-representation of genders in the candidate slate can all impact the final decision. By decoupling sources of bias, we can better isolate strategies for bias mitigation in human-in-the-loop systems.
研究动机与目标
- 隔离并测量人类招聘决策偏见的影响,独立于世界分布和算法偏见。
- 评估代表标准——候选人名单中固定比例的性别展示——是否能减轻招聘决策中的人类性别偏见。
- 理解决策者性别、任务复杂性以及过度或不足代表等因素如何影响招聘结果。
- 通过解耦世界、算法和人类贡献,分解混合型人机协同招聘系统中的偏见来源。
- 提供实证证据,以指导设计更公平的招聘系统,明确代表标准在何时有效或不足。
提出的方法
- 通过亚马逊机械 Turk 开展大规模受控实验,模拟多样化职业的招聘决策。
- 生成具有相同资历但随机化性别和代词的候选人档案,同时在每个名单中保持固定的性别比例。
- 通过确保候选人名单具有预设性别分布(例如 50-50、3-97)来应用代表标准,以隔离代表性的效应。
- 通过要求参与者从 8 名候选人中选出其前 4 名,测量人类推荐结果,模拟推荐任务。
- 与基于真实世界性别分布的基线以及使用词嵌入训练的 AI 模型进行对比,评估无干预情况下的偏见。
- 按决策者性别和职业分析数据,以评估对决策的混杂影响。
实验结果
研究问题
- RQ1在候选人名单中平衡性别代表性是否能减轻招聘决策中的人类性别偏见?
- RQ2代表标准在不同世界性别分布的职业中的有效性如何变化?
- RQ3在代表标准无法减少偏见的职业中,对代表性不足性别的过度代表是否能改善结果?
- RQ4决策者个人特征(尤其是其性别)如何影响招聘推荐?
- RQ5任务复杂性以及候选人名单构成(如过度或不足代表)如何影响人类在招聘中的决策?
主要发现
- 在世界分布严重失衡的职业中,如软件工程和法律,代表标准显著减少了性别偏见。
- 在存在持续人类偏好的职业中,如管道工或建筑行业,即使名单中的性别分布已平衡,代表标准也无法纠正偏见。
- 决策者性别影响结果:在某些职业中,女性决策者表现出的偏见少于男性决策者,但影响程度各异。
- 任务复杂性和候选人名单中代表性不足或过度代表的程度会干扰招聘决策,其中代表性不足会加剧偏见。
- 在以男性为主导的领域中,过度代表女性(如名单中 50% 为女性)并不能一致地改善公平性结果,尤其是在根深蒂固的偏好持续存在的情况下。
- 本研究表明,代表标准并非万能方案;其有效性取决于世界分布、人类偏好与决策者特征之间的相互作用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。