Skip to main content
QUICK REVIEW

[论文解读] Bounding and Approximating Intersectional Fairness through Marginal Fairness

Mathieu Molina, Patrick Loiseau|arXiv (Cornell University)|Jun 12, 2022
Ethics and Social Impacts of AI被引用 4
一句话总结

本文提出了一种统计框架,通过边缘公平度量来界定和近似机器学习中的交叉公平性。通过将受保护属性建模为随机变量,该框架利用边缘分布和依赖度量推导出交叉不公平性的高概率界,并提出一种启发式方法,通过基于统计独立性的属性分组来改进界,该方法在真实和合成数据集上得到验证。

ABSTRACT

Discrimination in machine learning often arises along multiple dimensions (a.k.a. protected attributes); it is then desirable to ensure \emph{intersectional fairness} -- i.e., that no subgroup is discriminated against. It is known that ensuring \emph{marginal fairness} for every dimension independently is not sufficient in general. Due to the exponential number of subgroups, however, directly measuring intersectional fairness from data is impossible. In this paper, our primary goal is to understand in detail the relationship between marginal and intersectional fairness through statistical analysis. We first identify a set of sufficient conditions under which an exact relationship can be obtained. Then, we prove bounds (easily computable through marginal fairness and other meaningful statistical quantities) in high-probability on intersectional fairness in the general case. Beyond their descriptive value, we show that these theoretical bounds can be leveraged to derive a heuristic improving the approximation and bounds of intersectional fairness by choosing, in a relevant manner, protected attributes for which we describe intersectional subgroups. Finally, we test the performance of our approximations and bounds on real and synthetic data-sets.

研究动机与目标

  • 理解在存在多个受保护属性的情况下,边缘公平性与交叉公平性之间的关系。
  • 仅使用边缘公平性和统计依赖度量,推导出可计算的、高概率的交叉不公平性界。
  • 提出一种启发式方法,对受保护属性进行分组,以提高公平性界和近似值的紧致性。
  • 在真实和合成数据集上验证理论界和启发式方法,证明其在数据稀缺情形下的实际效用。
  • 提供一种适用于标准公平性概念(如人口均等和机会均等)的一般性框架。

提出的方法

  • 将受保护属性建模为离散型随机变量,并使用总相关性来量化它们之间的统计依赖性。
  • 利用边缘公平度量和受保护属性的总相关性,推导出交叉不公平性的概率界。
  • 提出一种划分启发式方法,将具有高统计依赖性的属性分组,以减少方差并收紧界。
  • 将界改进表述为对数似然比方差和依赖结构的函数,通过使用更粗粒度的划分来减少总相关性。
  • 通过经验估计边缘分布和依赖度量,在实践中计算界和近似值。
  • 通过在真实(1990年美国人口普查)和合成数据集上的实验验证该方法,比较界紧致性和收敛速度。

实验结果

研究问题

  • RQ1在何种条件下,可以仅从边缘公平性度量中精确恢复交叉公平性?
  • RQ2如何仅使用边缘公平性和统计依赖信息来界定交叉不公平性?
  • RQ3能否基于统计依赖性对受保护属性进行分组,以提高公平性界的紧致性?
  • RQ4在高维稀疏数据情形下,所提出的界与交叉不公平性的经验估计相比如何?
  • RQ5数据稀疏性对边缘公平性作为交叉公平性代理的可靠性有何影响?

主要发现

  • 本文识别出在何种充分条件下,可从边缘公平性中精确推导出交叉公平性,明确了边缘公平性作为代理的局限性。
  • 利用边缘公平性和总相关性,推导出交叉不公平性的高概率界,这些度量可轻松从数据中估计。
  • 所提出的基于统计独立性的受保护属性分组启发式方法,在所有数据集中一致地提高了界紧致性。
  • 实证结果表明,改进后的界比基线界更紧致,并且更接近真实交叉不公平性,尤其是在稀疏数据情形下。
  • 即使每个子群的数据量有限,该方法依然有效,当子群规模太小而无法可靠进行经验测量时,其表现优于直接估计。
  • 当由于子群数量呈指数增长而导致精确交叉公平性估计不可行时,该界被证明在实践中具有实用性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。