Skip to main content
QUICK REVIEW

[Paper Review] Learning Optimal Fair Policies

Razieh Nabi, Daniel Malinsky|arXiv (Cornell University)|Sep 6, 2018
Experimental Behavioral Economics Studies24 references22 citations
TL;DR

This paper proposes a causal inference and constrained optimization framework to learn optimal decision policies that are fair with respect to sensitive attributes (e.g., race, gender), ensuring fairness in both decisions and outcomes. It formally constrains unfair causal pathways from sensitive features to actions and outcomes, guaranteeing that the induced joint distribution satisfies fairness constraints while maximizing expected utility, thus 'breaking the cycle of injustice' in data-driven policy making.

ABSTRACT

Systematic discriminatory biases present in our society influence the way data is collected and stored, the way variables are defined, and the way scientific findings are put into practice as policy. Automated decision procedures and learning algorithms applied to such data may serve to perpetuate existing injustice or unfairness in our society. In this paper, we consider how to make optimal but fair decisions, which "break the cycle of injustice" by correcting for the unfair dependence of both decisions and outcomes on sensitive features (e.g., variables that correspond to gender, race, disability, or other protected attributes). We use methods from causal inference and constrained optimization to learn optimal policies in a way that addresses multiple potential biases which afflict data analysis in sensitive contexts, extending the approach of (Nabi and Shpitser 2018). Our proposal comes equipped with the theoretical guarantee that the chosen fair policy will induce a joint distribution for new instances that satisfies given fairness constraints. We illustrate our approach with both synthetic data and real criminal justice data.

Motivation & Objective

  • To address systemic biases in data-driven policy learning that perpetuate unfair outcomes based on sensitive attributes such as race or gender.
  • To formalize fairness as constraints on specific causal pathways from sensitive features to decisions and outcomes.
  • To develop a method that learns optimal policies under fairness constraints while maintaining high utility and theoretical guarantees on fairness in the induced joint distribution.
  • To demonstrate the method’s effectiveness on synthetic data and real-world criminal justice data with a heuristic utility function.
  • To show that fair policy learning can reduce racial disparities in incarceration rates without sacrificing overall policy quality.

Proposed method

  • Uses causal inference to model the impact of sensitive features on decisions and outcomes through structural causal models.
  • Applies mediation analysis to identify and constrain direct and indirect causal pathways from sensitive attributes (S) to actions (A) and outcomes (Y).
  • Imposes fairness constraints on the S→A and S→Y pathways to prevent unfair influence of sensitive features.
  • Employs constrained optimization to learn a policy that maximizes expected utility while satisfying fairness constraints.
  • Uses Q-learning with a parametric utility function Y = (1−A)(−θR + (1−R)) − A to optimize expected utility under fairness constraints.
  • Applies semiparametric estimation techniques to ensure robustness and theoretical guarantees under model misspecification.

Experimental results

Research questions

  • RQ1How can we learn optimal policies that are fair with respect to sensitive attributes while maintaining high utility?
  • RQ2What causal pathways from sensitive features to decisions and outcomes must be constrained to ensure fairness?
  • RQ3Can we guarantee that the joint distribution induced by a fair policy satisfies specified fairness constraints?
  • RQ4How does the choice of utility function affect the trade-off between fairness and policy performance in real-world settings?
  • RQ5To what extent can fair policy learning reduce racial disparities in incarceration rates in criminal justice systems?

Key findings

  • The proposed method successfully learns policies that induce a joint distribution satisfying fairness constraints, effectively 'breaking the cycle of injustice' in decision-making.
  • For θ between 2 and 3, the fair policy reduces overall incarceration rates compared to observed rates and narrows the racial gap in incarceration, though does not eliminate it entirely.
  • When θ > 3, both fair and unconstrained policies recommend higher-than-observed incarceration rates, but the fair policy still achieves a narrower racial gap.
  • The fair policy consistently reduces racial disparities in incarceration rates across a range of utility parameter values, demonstrating its effectiveness in mitigating bias.
  • The method performs well on real criminal justice data, showing that fairness constraints can be practically implemented without sacrificing policy quality.
  • The results illustrate that the choice of utility function significantly influences policy outcomes, highlighting the importance of careful utility specification in fairness-aware policy learning.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.