[论文解读] Fairness in Risk Assessment Instruments: Post-Processing to Achieve Counterfactual Equalized Odds
本文提出一种后处理方法,通过双重稳健估计器调整现有预测器,实现在风险评估工具(RAIs)中的反事实等 odds 公平性。该方法通过针对潜在结果而非可观测结果来确保公平性,并以快速收敛速率逼近最优公平预测器,在高风险决策场景中优于基于观测结果的公平性标准。
In domains such as criminal justice, medicine, and social welfare, decision makers increasingly have access to algorithmic Risk Assessment Instruments (RAIs). RAIs estimate the risk of an adverse outcome such as recidivism or child neglect, potentially informing high-stakes decisions such as whether to release a defendant on bail or initiate a child welfare investigation. It is important to ensure that RAIs are fair, so that the benefits and harms of such decisions are equitably distributed. The most widely used algorithmic fairness criteria are formulated with respect to observable outcomes, such as whether a person actually recidivates, but these criteria are misleading when applied to RAIs. Since RAIs are intended to inform interventions that can reduce risk, the prediction itself affects the downstream outcome. Recent work has argued that fairness criteria for RAIs should instead utilize potential outcomes, i.e. the outcomes that would occur in the absence of an appropriate intervention. However, no methods currently exist to satisfy such fairness criteria. In this paper, we target one such criterion, counterfactual equalized odds. We develop a post-processed predictor that is estimated via doubly robust estimators, extending and adapting previous post-processing approaches to the counterfactual setting. We also provide doubly robust estimators of the risk and fairness properties of arbitrary fixed post-processed predictors. Our predictor converges to an optimal fair predictor at fast rates. We illustrate properties of our method and show that it performs well on both simulated and real data.
研究动机与目标
- 解决基于观测结果的观测公平性标准在 RAIs 中的局限性,这些标准依赖于可观测结果,当干预措施影响风险时可能产生误导。
- 提出一种新的公平性准则——近似反事实等 odds,该准则考虑在无干预条件下的潜在结果,更真实地反映风险预测的初衷。
- 开发一种后处理框架,将任意现有预测器转换为满足近似反事实等 odds 的预测器。
- 提供后处理预测器收敛速率的理论保证,其依赖于干扰参数估计的准确性。
- 在模拟数据和真实世界数据(包括儿童福利和刑事司法数据集)上展示方法的实证性能。
提出的方法
- 该方法将近似反事实等 odds(cEO)定义为基于无治疗条件下潜在结果的公平性准则。
- 通过线性规划计算损失最优的后处理预测器,以满足 cEO 约束条件。
- 使用双重稳健估计器来估计干扰参数(如结果模型和倾向得分模型),确保只要其中任一模型正确指定,估计结果即具有一致性。
- 后处理预测器在运行时仅依赖于敏感特征和原始预测器,便于在现有 RAIs 上轻松部署。
- 该方法将先前的后处理方法(如 Hardt 等,2016 年)扩展至反事实设置,其中公平性定义在潜在结果之上。
- 理论分析表明,预测器以快速收敛速率逼近最优公平预测器,其收敛速度取决于干扰估计器的收敛性。

实验结果
研究问题
- RQ1是否可以将 RAIs 的公平性准则重新定义为基于潜在结果而非可观测结果,以更真实地反映风险预测的真正目的?
- RQ2是否可能开发一种后处理方法,在无需重新训练或内嵌处理的情况下强制实现反事实等 odds 公平性?
- RQ3在双重稳健框架中,后处理预测器的收敛速率如何依赖于干扰模型估计的质量?
- RQ4在真实世界决策系统中,该方法是否能有效降低歧视性差异影响,优于观测公平性标准?
- RQ5在反事实等 odds 公平性准则下,公平性与预测性能之间存在何种权衡?
主要发现
- 所提出的后处理预测器通过基于双重稳健估计的潜在结果调整预测,实现了近似反事实等 odds 公平性。
- 该方法以快速收敛速率逼近最优公平预测器,收敛速度取决于干扰参数估计的准确性。
- 在模拟实验中,该方法在中等误报与误报成本比率下有效平衡了公平性与性能。
- 在儿童福利服务的真实数据上,该方法在整体准确率未显著下降的情况下产生了更公平的预测结果。
- 当错误成本比率变得极端时,后处理分类器才会趋近于平凡分类器(始终输出 0 或始终输出 1),表明对输入敏感性具有鲁棒性。
- 该方法可在现有 RAIs 上可行部署,仅需在推理时访问敏感特征和原始预测器。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。