Skip to main content
QUICK REVIEW

[论文解读] Fair Regression with Wasserstein Barycenters

Evgenii Chzhen, Christophe Denis|HAL (Le Centre pour la Communication Scientifique Directe)|Jun 12, 2020
Health Systems, Economic Evaluations, Quality of Life被引用 17
一句话总结

本文提出了一种公平回归方法,通过将群体特定回归分布的Wasserstein均值作为最优预测器,确保了人口均等性。该方法引入了一种后处理算法,可将任意现成的回归器转化为公平回归器,在无需重新训练的情况下,实现了无分布公平性保证和有限样本风险保证,且误差增长优于公平性增益。

ABSTRACT

We study the problem of learning a real-valued function that satisfies the Demographic Parity constraint. It demands the distribution of the predicted output to be independent of the sensitive attribute. We consider the case that the sensitive attribute is available for prediction. We establish a connection between fair regression and optimal transport theory, based on which we derive a close form expression for the optimal fair predictor. Specifically, we show that the distribution of this optimum is the Wasserstein barycenter of the distributions induced by the standard regression function on the sensitive groups. This result offers an intuitive interpretation of the optimal fair prediction and suggests a simple post-processing algorithm to achieve fairness. We establish risk and distribution-free fairness guarantees for this procedure. Numerical experiments indicate that our method is very effective in learning fair models, with a relative increase in error rate that is inferior to the relative gain in fairness.

研究动机与目标

  • 解决在人口均等性约束下学习公平实值预测器的挑战。
  • 通过Wasserstein均值建立公平回归与最优传输理论之间的理论联系。
  • 开发一种后处理方法,确保公平性独立于基础回归器和数据分布。
  • 为所提方法提供有限样本风险界和无分布公平性保证。

提出的方法

  • 最优公平预测器被推导为不公平回归函数在每个敏感群体上诱导的条件分布的Wasserstein均值。
  • 该方法使用原始回归函数的闭式变换:$ g^*(x,s) = p_s f^*(x,s) + (1-p_s) t^*(x,s) $,其中 $ t^* $ 使各群体间的排名对齐。
  • 校正项 $ t^* $ 通过各群体的样本累积分布函数计算,仅依赖于各群体的无标签数据。
  • 提出了一种后处理程序,在估计基础回归函数后应用此变换,确保公平性而无需重新训练。
  • 该方法利用一维Wasserstein距离的秩统计量和偏差界,推导出有限样本风险保证。
  • 理论分析结合了最优传输、秩统计量和浓度不等式的结果,建立了收敛速率。

实验结果

研究问题

  • RQ1如何利用最优传输理论在回归任务中正式强制实现人口均等性?
  • RQ2在人口均等性约束下,最优公平预测器的解析形式是什么?
  • RQ3能否设计一种后处理方法,确保公平性独立于基础估计器或底层数据分布?
  • RQ4在最小假设下,能否为公平预测器推导出有限样本风险界?
  • RQ5在实践中,公平预测器的误差与公平性增益相比如何?

主要发现

  • 最优公平预测器对应于群体特定回归分布的Wasserstein均值,为公平性提供了几何解释。
  • 所提后处理方法实现了无分布公平性,即无论数据分布或基础估计器如何,公平性均能保证。
  • 建立了有限样本风险界,表明公平预测器的期望误差以 $ O(b_n^{-1/2} + extstyle rac{1}{N^{1/2}}) $ 的速率收敛,其中 $ b_n $ 为带宽,$ N $ 为样本量。
  • 数值实验确认,预测误差的相对增加始终低于公平性增益的相对提升。
  • 该方法具有鲁棒性和可扩展性,仅需无标签数据完成变换步骤,且可应用于任意现成的回归模型。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。