Skip to main content
QUICK REVIEW

[論文レビュー] To the Fairness Frontier and Beyond: Identifying, Quantifying, and Optimizing the Fairness-Accuracy Pareto Frontier

Camille Olivia Little, Michael Weylandt|arXiv (Cornell University)|May 31, 2022
Ethics and Social Impacts of AI被引用数 4
ひとこと要約

本稿では、任意のグループ公平性定義および精度測定法に対して、公平性と精度のパレート前線を経験的に特徴付け、定量化するための taf カーブと、Fairness-Area-Under-the-Curve (fauc) メトリックを導入する。さらに、凸最適化に基づくモデルスタッキングフレームワークである FairStacks を提案し、個々のモデルの限界を超えて経験的パレート前線を拡張し、公平性制約の下で精度を最大化することで fauc を向上させる。このフレームワークはベンチマークデータセットにおいて、既存手法を上回る性能を示した。

ABSTRACT

Algorithmic fairness has emerged as an important consideration when using machine learning to make high-stakes societal decisions. Yet, improved fairness often comes at the expense of model accuracy. While aspects of the fairness-accuracy tradeoff have been studied, most work reports the fairness and accuracy of various models separately; this makes model comparisons nearly impossible without a model-agnostic metric that reflects the balance of the two desiderata. We seek to identify, quantify, and optimize the empirical Pareto frontier of the fairness-accuracy tradeoff. Specifically, we identify and outline the empirical Pareto frontier through Tradeoff-between-Fairness-and-Accuracy (TAF) Curves; we then develop a metric to quantify this Pareto frontier through the weighted area under the TAF Curve which we term the Fairness-Area-Under-the-Curve (FAUC). TAF Curves provide the first empirical, model-agnostic characterization of the Pareto frontier, while FAUC provides the first metric to impartially compare model families on both fairness and accuracy. Both TAF Curves and FAUC can be employed with all group fairness definitions and accuracy measures. Next, we ask: Is it possible to expand the empirical Pareto frontier and thus improve the FAUC for a given collection of fitted models? We answer affirmately by developing a novel fair model stacking framework, FairStacks, that solves a convex program to maximize the accuracy of model ensemble subject to a score-bias constraint. We show that optimizing with FairStacks always expands the empirical Pareto frontier and improves the FAUC; we additionally study other theoretical properties of our proposed approach. Finally, we empirically validate TAF, FAUC, and FairStacks through studies on several real benchmark data sets, showing that FairStacks leads to major improvements in FAUC that outperform existing algorithmic fairness approaches.

研究の動機と目的

  • 多様なモデルファミリーにわたる、経験的公平性-精度パレート前線の特定、定量的評価、最適化を目的とする。
  • 公平性と精度の両面を一貫して比較可能な、モデルに依存しない統一メトリックの開発を目的とする。
  • 高リスクな機械学習意思決定における公平性と精度のバランスという重要な課題に取り組むことを目的とする。
  • 個々のモデルの限界を超えて経験的パレート前線を拡張するメタラーニングフレームワークの提案を目的とする。
  • 提案されたフレームワークを実世界のベンチマークデータセットに対して経験的に検証することを目的とする。

提案手法

  • 任意のモデルに対して公平性-精度パレート前線のモデルに依存しない経験的特徴付けとして taf カーブを提案し、各公平性レベルにおける最大精度をプロットする。
  • 公平性-精度トレードオフ全体を定量的に評価するための、taf カーブの下側に位置する重み付き面積として Fairness-Area-Under-the-Curve (fauc) メトリックを導入する。
  • スコアベースの公平性制約の下で精度を最大化するように、凸最適化問題を解くことでモデルを組み合わせる公平なモデルスタッキングフレームワーク FairStacks を開発する。
  • 事前に学習済みモデルの線形結合を制約付き凸計画問題で最適化し、あらゆる公平性レベルで公平性を確保しながら精度を向上させる。
  • 実世界のデータセットにフレームワークを適用し、経験的パレート前線および fauc の一貫した向上を実証した。
  • モデルに依存しない設計により、すべてのグループ公平性定義および精度測定法への一般化を可能にした。

実験結果

リサーチクエスチョン

  • RQ1任意の適合済みモデルの集合に対して、モデルに依存しない方法で公平性-精度パレート前線を経験的に特徴付けられるか?
  • RQ2異なるモデルや公平性定義にわたって、公平性-精度トレードオフを統一的かつ解釈可能な形で定量的に評価するメトリックは存在するか?
  • RQ3メタラーニングを用いることで、個々のモデルの性能を超えて経験的パレート前線を拡張できるか?
  • RQ4FairStacks フレームワークは、多様なデータセットおよび公平性基準において、一貫して公平性と精度のトレードオフを改善できるか?
  • RQ5提案されたフレームワークは、公平性および精度の両面において、既存のアルゴリズム的手法と比較して優れているか?

主な発見

  • taf カーブは、任意の公平性定義および精度測定法に対して、公平性-精度パレート前線を経験的かつモデルに依存しない形で可視化する初の手法である。
  • fauc メトリックは、公平性-精度のバランスに基づいてモデルファミリーを定量的に比較するための単一の統一スコアを提供する。
  • FairStacks は、個々のモデルよりもあらゆる公平性レベルで高い精度を達成することで、経験的パレート前線を一貫して拡張する。
  • 複数のベンチマークデータセットにおいて、fauc スコアを顕著に向上させ、既存の公平性緩和手法を上回る性能を示した。
  • FairStacks を用いた最適化は、公平性制約の下で精度を最大化する凸計画問題を解くため、パレート前線が保証的に拡張される。
  • 本手法は汎用的であり、任意のグループ公平性定義および精度測定法に適用可能であり、広範な実用的展開が可能である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。