Skip to main content
QUICK REVIEW

[论文解读] Omitted and Included Variable Bias in Tests for Disparate Impact

Jongbin Jung, Sam Corbett‐Davies|arXiv (Cornell University)|Sep 15, 2018
Advanced Causal Inference Techniques参考文献 30被引用 6
一句话总结

本文提出一种三步风险调整回归方法,以解决在差异影响检测中因遗漏变量和包含无关变量导致的偏差问题。首先,利用所有协变量估计决策效用,形成风险评分;其次,仅以这些估计值作为控制变量,测试群体间的差异;最后,评估未测量混杂因素对结果的影响。该方法通过减少歧视检测中的偏差,提升了结果的可靠性——在220万起纽约市行人盘查数据中得到验证,传统方法在此类数据中曾产生误导性结论。

ABSTRACT

Policymakers often seek to gauge discrimination against groups defined by race, gender, and other protected attributes. One popular strategy is to estimate disparities after controlling for observed covariates, typically with a regression model. This approach, however, suffers from two statistical challenges. First, omitted-variable bias can skew results if the model does not control for all relevant factors; second, and conversely, included-variable bias can skew results if the set of controls includes irrelevant factors. Here we introduce a simple three-step strategy---which we call risk-adjusted regression---that addresses both concerns in settings where decision makers have clearly measurable objectives. In the first step, we use all available covariates to estimate the utility of possible decisions. In the second step, we measure disparities after controlling for these utility estimates alone, mitigating the problem of included-variable bias. Finally, in the third step, we examine the sensitivity of results to unmeasured confounding, addressing concerns about omitted-variable bias. We demonstrate this method on a detailed dataset of 2.2 million police stops of pedestrians in New York City, and show that traditional statistical tests of discrimination can yield misleading results. We conclude by discussing implications of our statistical approach for questions of law and policy.

研究动机与目标

  • 解决统计检验中因遗漏变量导致的偏差问题,即未测量的混杂因素会扭曲歧视检测结果。
  • 纠正因包含无关或冗余控制变量而引起的包含变量偏差,此类变量会扭曲群体差异估计。
  • 开发一种方法,使统计检验与政策相关场景中可明确衡量的决策目标保持一致。
  • 提供针对未测量混杂因素的敏感性分析框架,提升歧视发现结果的稳健性。
  • 通过220万起纽约市行人盘查的大规模数据集,展示该方法的实际影响。

提出的方法

  • 步骤1:利用所有可用协变量预测结果,估计决策效用,形成风险评分。
  • 步骤2:通过仅控制估计出的风险评分,回归分析结果与群体归属关系,以减少包含变量偏差。
  • 步骤3:通过改变对未测量混杂因素的假设,进行敏感性分析,以评估其对结果的影响,从而应对遗漏变量偏差。
  • 该方法依赖于完整协变量模型生成的风险评分,将其作为公平性检验的充分统计量。
  • 利用回归模型在控制效用估计后估计差异,确保仅相关因素影响结果。
  • 通过调整未测量混杂因素的假设,执行敏感性分析,以检验发现结果的稳健性。

实验结果

研究问题

  • RQ1在现实政策应用中,遗漏变量偏差如何扭曲差异影响的统计检验?
  • RQ2在歧视检测中,包含无关或冗余协变量在多大程度上引入偏差?
  • RQ3风险调整回归框架能否同时缓解遗漏变量和包含变量偏差?
  • RQ4传统歧视检验在大规模执法数据中对未测量混杂因素的敏感性如何?
  • RQ5在考虑决策效用和未测量混杂因素的情况下,使用该方法的政策含义是什么?

主要发现

  • 在执法数据中,传统歧视检验方法可能因未控制的混杂因素和无关协变量而产生误导性结果。
  • 风险调整回归方法通过聚焦于基于所有协变量估计的效用值,减少了无关控制变量的影响,从而降低了偏差。
  • 敏感性分析显示,未测量混杂因素可能显著改变关于差异影响的结论,凸显了稳健性检验的必要性。
  • 在纽约市行人盘查数据集中,该方法识别出了在标准回归方法下被掩盖或反转的差异。
  • 该方法通过将统计建模与决策目标对齐,为法律和政策评估提供了更具说服力的框架。
  • 本研究证明,风险调整回归能提升高风险、现实场景中公平性评估的有效性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。