Skip to main content
QUICK REVIEW

[论文解读] Learning Optimal Fair Policies

Razieh Nabi, Daniel Malinsky|arXiv (Cornell University)|Sep 6, 2018
Experimental Behavioral Economics Studies参考文献 24被引用 22
一句话总结

本文提出了一种因果推断与约束优化框架,用于学习在敏感属性(如种族、性别)方面公平的最优决策策略,确保决策与结果均公平。该方法正式约束了从敏感特征到行动和结果的不公平因果路径,保证所诱导的联合分布满足公平约束,同时最大化期望效用,从而‘打破数据驱动政策制定中的不公循环’。

ABSTRACT

Systematic discriminatory biases present in our society influence the way data is collected and stored, the way variables are defined, and the way scientific findings are put into practice as policy. Automated decision procedures and learning algorithms applied to such data may serve to perpetuate existing injustice or unfairness in our society. In this paper, we consider how to make optimal but fair decisions, which "break the cycle of injustice" by correcting for the unfair dependence of both decisions and outcomes on sensitive features (e.g., variables that correspond to gender, race, disability, or other protected attributes). We use methods from causal inference and constrained optimization to learn optimal policies in a way that addresses multiple potential biases which afflict data analysis in sensitive contexts, extending the approach of (Nabi and Shpitser 2018). Our proposal comes equipped with the theoretical guarantee that the chosen fair policy will induce a joint distribution for new instances that satisfies given fairness constraints. We illustrate our approach with both synthetic data and real criminal justice data.

研究动机与目标

  • 解决数据驱动政策学习中因种族或性别等敏感属性而持续产生不公平结果的系统性偏见问题。
  • 将公平性形式化为从敏感特征到决策与结果的特定因果路径上的约束。
  • 开发一种在公平约束下学习最优策略的方法,同时保持高效率,并对所诱导的联合分布提供公平性的理论保证。
  • 在合成数据和现实世界刑事司法数据上,通过启发式效用函数展示该方法的有效性。
  • 表明公平策略学习可在不牺牲整体政策质量的前提下,减少刑事司法系统中的种族差异。

提出的方法

  • 使用因果推断,通过结构因果模型来建模敏感特征对决策与结果的影响。
  • 应用中介分析,识别并约束从敏感属性(S)到行动(A)和结果(Y)的直接与间接因果路径。
  • 对 S→A 和 S→Y 路径施加公平性约束,以防止敏感特征产生不公平影响。
  • 采用约束优化方法,学习在满足公平约束下最大化期望效用的策略。
  • 使用 Q-learning 与参数化效用函数 Y = (1−A)(−θR + (1−R)) − A,在公平约束下优化期望效用。
  • 采用半参数估计技术,以在模型误设情况下确保稳健性与理论保证。

实验结果

研究问题

  • RQ1如何在保持高效率的同时,学习对敏感属性公平的最优策略?
  • RQ2从敏感特征到决策与结果的哪些因果路径必须被约束,才能确保公平性?
  • RQ3能否保证由公平策略诱导的联合分布满足指定的公平约束?
  • RQ4效用函数的选择如何影响现实场景中公平性与政策表现之间的权衡?
  • RQ5公平策略学习在多大程度上可以减少刑事司法系统中的种族差异?

主要发现

  • 所提出的方法成功学习到能诱导满足公平约束的联合分布的策略,有效‘打破决策中的不公循环’。
  • 当 θ 在 2 到 3 之间时,公平策略相比观察到的比率降低了总体监禁率,并缩小了种族间的监禁差距,尽管并未完全消除。
  • 当 θ > 3 时,公平策略与无约束策略均建议高于观察到的监禁率,但公平策略仍实现了更小的种族差距。
  • 在一系列效用参数值下,公平策略始终减少了监禁率中的种族差异,证明其在减轻偏见方面的有效性。
  • 该方法在真实刑事司法数据上表现良好,表明公平约束可实际实施且不牺牲政策质量。
  • 结果表明,效用函数的选择显著影响政策结果,强调了在公平感知策略学习中仔细指定效用函数的重要性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。