Skip to main content
QUICK REVIEW

[論文レビュー] Entrofy Your Cohort: A Data Science Approach to Candidate Selection

Daniela Huppenkothen, Brian McFee|arXiv (Cornell University)|May 8, 2019
Names, Identity, and Discrimination Research参考文献 36被引用数 4
ひとこと要約

本論文では、盲検された評価後、多様性志向のコhort選抜を自動化するデータサイエンスアルゴリズムであるEntrofyを紹介する。所定の人口統計的およびカテゴリカル基準を最適化することで、透明性があり、監査可能で、バイアスに強くない高評価候補者の選抜を保証する。シミュレーションおよびAstro Hack Week 2016の事例研究により検証済み。

ABSTRACT

Selecting a cohort from a set of candidates is a common task within and beyond academia. Admitting students, awarding grants, choosing speakers for a conference are situations where human biases may affect the make-up of the final cohort. We propose a new algorithm, Entrofy, designed to be part of a larger decision making strategy aimed at making cohort selection as just, quantitative, transparent, and accountable as possible. We suggest this algorithm be embedded in a two-step selection procedure. First, all application materials are stripped of markers of identity that could induce conscious or sub-conscious bias. During blind review, the committee selects all applicants, submissions, or other entities that meet their merit-based criteria. This often yields a cohort larger than the admissible number. In the second stage, the target cohort can be chosen from this meritorious pool via a new algorithm and software tool. Entrofy optimizes differences across an assignable set of categories selected by the human committee. Criteria could include gender, academic discipline, experience with certain technologies, or other quantifiable characteristics. The Entrofy algorithm yields the computational maximization of diversity by solving the tie-breaking problem with provable performance guarantees. We show how Entrofy selects cohorts according to pre-determined characteristics in simulated sets of applications and demonstrate its use in a case study. This cohort selection process allows human judgment to prevail when assessing merit, but assigns the assessment of diversity to a computational process less likely to be beset by human bias. Importantly, the stage at which diversity assessments occur is fully transparent and auditable with Entrofy. Splitting merit and diversity considerations into their own assessment stages makes it easier to explain why a given candidate was selected or rejected.

研究の動機と目的

  • 特に入学選抜、助成金交付、国際会議のスピーカー選定において、人間のバイアスが生じる学術的および職業的コhort選抜プロセスに対処すること。
  • 評価の妥当性と多様性の考慮を分離することで、内的バイアスを低減し、公平性を向上させること。
  • 高評価候補者からなるプールから、透明性があり、監査可能で、再現可能な多様なコhortを選び出すための計算ツールを開発すること。
  • 委員会の意思決定で一般的に使われる直感的またはヒューリスティックな方法の代替として、スケーラブルで定量的な代替案を提供すること。
  • ワークショップや会議などの現実の学術的環境において、アルゴリズム的多様性最大化の実現可能性と有効性を示すこと。

提案手法

  • 2段階の選抜プロセスを実装する:まず、名前、所属機関、アイデンティティの兆候を削除することで、全候補者の資格を盲検された形で評価する。
  • 2番目に、ユーザーが定義した多様性基準に基づき、評価に合格した候補者プールから最終的なコhortをEntrofyアルゴリズムで選抜する。
  • 性別、国、専門分野、技術的経験などの複数のカテゴリカル属性における多様性を最大化する最適化フレームワークを用い、保証された性能を達成する。
  • 多様性の目標と高い評価を維持する必要という両立を図る制約付き最適化問題として選抜を定式化する。
  • 組合せ最適化技術を活用し、数学的に整合性があり、効率的な方法で同点の解を決定する。
  • Entrofyアルゴリズムで使用されたすべての意思決定と基準をログに記録することで、プロセス全体の透明性と監査可能性を確保する。

実験結果

リサーチクエスチョン

  • RQ1従来の委員会ベースの方法と比較して、アルゴリズム的アプローチは、コhort選抜における公平性と透明性を向上させることができるか?
  • RQ2評価の妥当性を保ちながら、どの程度まで多様性を体系的に最大化できるか?
  • RQ3盲検された評価の後、アルゴリズム的多様性選抜を実施する2段階プロセスは、最終コhortの構成と責任の所在にどのような影響を与えるか?
  • RQ4複数のカテゴリカル制約がある現実のシナリオにおいて、Entrofyアルゴリズムは、最適または近似最適な多様性結果を信頼性を持って得られるか?
  • RQ5多様な候補者プールを有する学術的ワークショップ、会議、入学選抜プロセスにおいて、Entrofyを使用する実用的意義は何か?

主な発見

  • Entrofyは、シミュレーテッドデータセットおよび実世界の応用例(例:Astro Hack Week 2016)において、所定の多様性目標を達成したコhortを効果的に生成した。
  • テストされたすべてのシナリオで、最適または近似最適の多様性スコアを達成し、最適化フレームワーク下での強い性能保証を示した。
  • 2段階プロセス(盲検された評価 → アルゴリズム的多様性選抜)は、従来の委員会ベースの選抜よりも、より透明性があり、監査可能な結果をもたらした。
  • 評価の妥当性と多様性の評価を分離することで、最終選抜における内的バイアスのリスクが低減された。
  • Astro Hack Week 2016の事例研究では、性別、キャリアステージ、地理的起源の面でバランスの取れたコhortを、高い評価基準を維持したまま生成できた。
  • オープンソースのソフトウェアと完全に再現可能なコード(GitHubで公開)のおかげで、学術的および職業的環境での再現と導入が可能になった。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。