Skip to main content
QUICK REVIEW

[论文解读] Entrofy Your Cohort: A Data Science Approach to Candidate Selection

Daniela Huppenkothen, Brian McFee|arXiv (Cornell University)|May 8, 2019
Names, Identity, and Discrimination Research参考文献 36被引用 4
一句话总结

本文介绍了 Entrofy,一种数据科学算法,可在盲评 merit 审查后自动实现以多样性为导向的队列选择。通过优化预设的人口统计和类别标准,Entrofy 确保了高 merit 候选人的透明、可审计且抗偏见的选拔,该方法在模拟和 2016 年 Astro Hack Week 的案例研究中得到验证。

ABSTRACT

Selecting a cohort from a set of candidates is a common task within and beyond academia. Admitting students, awarding grants, choosing speakers for a conference are situations where human biases may affect the make-up of the final cohort. We propose a new algorithm, Entrofy, designed to be part of a larger decision making strategy aimed at making cohort selection as just, quantitative, transparent, and accountable as possible. We suggest this algorithm be embedded in a two-step selection procedure. First, all application materials are stripped of markers of identity that could induce conscious or sub-conscious bias. During blind review, the committee selects all applicants, submissions, or other entities that meet their merit-based criteria. This often yields a cohort larger than the admissible number. In the second stage, the target cohort can be chosen from this meritorious pool via a new algorithm and software tool. Entrofy optimizes differences across an assignable set of categories selected by the human committee. Criteria could include gender, academic discipline, experience with certain technologies, or other quantifiable characteristics. The Entrofy algorithm yields the computational maximization of diversity by solving the tie-breaking problem with provable performance guarantees. We show how Entrofy selects cohorts according to pre-determined characteristics in simulated sets of applications and demonstrate its use in a case study. This cohort selection process allows human judgment to prevail when assessing merit, but assigns the assessment of diversity to a computational process less likely to be beset by human bias. Importantly, the stage at which diversity assessments occur is fully transparent and auditable with Entrofy. Splitting merit and diversity considerations into their own assessment stages makes it easier to explain why a given candidate was selected or rejected.

研究动机与目标

  • 解决学术和职业队列选拔过程中的人类偏见,特别是在招生、资助和会议讲者选拔中的问题。
  • 将 merit 评估与多样性考量分离,以减少隐性偏见并提升公平性。
  • 开发一种计算工具,实现从高 merit 候选人池中透明、可审计且可复现的多样化队列选择。
  • 为委员会决策中常用的直观或基于启发式的方法提供可扩展、量化的替代方案。
  • 证明在现实学术场景(如研讨会和会议)中,算法化多样性最大化方法的可行性与有效性。

提出的方法

  • 实施两阶段选拔流程:首先进行盲评 merit 审查,以识别所有合格候选人,移除姓名、隶属关系和身份标识。
  • 其次,应用 Entrofy 算法,根据用户定义的多样性标准,从 merit 合格的候选人池中选择最终队列。
  • 采用一种优化框架,在多个类别属性(如性别、国家、学科、技术经验)上实现多样性最大化,并具备可证明的性能保障。
  • 将选拔问题形式化为一个约束优化问题,平衡多样性目标与保持高 merit 的要求。
  • 利用组合优化技术解决在数学上严谨且高效的方式下的选择冲突问题。
  • 通过记录 Entrofy 算法中使用的所有决策和标准,使整个过程保持透明和可审计。

实验结果

研究问题

  • RQ1与传统的委员会决策方法相比,算法化方法是否能提升队列选拔的公平性和透明度?
  • RQ2在保持基于 merit 的选拔的同时,多样性在多大程度上可以被系统性地最大化?
  • RQ3两阶段流程——即盲评 merit 审查后接算法化多样性选择——如何影响最终队列的构成与问责性?
  • RQ4在具有多个类别约束的现实场景中,Entrofy 算法是否能可靠地生成最优或近似最优的多样性结果?
  • RQ5在拥有多样化候选人池的学术研讨会、会议和招生流程中,使用 Entrofy 的实际影响是什么?

主要发现

  • Entrofy 在模拟数据集和真实应用场景(如 2016 年 Astro Hack Week)中,成功生成了符合预设多样性目标的队列。
  • 该算法在所有测试场景中均实现了最优或接近最优的多样性得分,证明了其在优化框架下的强大性能保障。
  • 两阶段流程——即盲评 merit 审查后接算法化多样性选择——相比传统委员会决策,产生了更透明和可审计的结果。
  • 该方法通过将 merit 评估与多样性评估解耦,降低了最终选拔中隐性偏见的风险。
  • 2016 年 Astro Hack Week 的案例研究显示,Entrofy 在性别、职业阶段和地理来源方面均能生成平衡的队列,同时保持了高 merit 标准。
  • 开源软件和完全可复现的代码(可在 GitHub 上获取)使得该方法可在学术和职业环境中广泛复制与采用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。