Skip to main content
QUICK REVIEW

[论文解读] Mitigating Included- and Omitted-Variable Bias in Estimates of Disparate Impact

Jongbin Jung, Sam Corbett‐Davies|arXiv (Cornell University)|Sep 15, 2018
Nuclear and radioactivity studies被引用 16
一句话总结

本文提出风险调整回归,一种三步法,以减轻在差异影响分析中因遗漏变量和包含无关变量导致的偏差。首先在不使用受保护属性的情况下估计风险,然后仅基于风险进行调整,最后对风险估计误差进行非参数敏感性分析。该方法揭示,传统回归方法低估了差异——例如,即使在控制了估计风险后,非裔和西班牙裔行人被搜身的概率仍显著高于白人,且其风险水平相当。

ABSTRACT

Managers, employers, policymakers, and others often seek to understand whether decisions are biased against certain groups. One popular analytic strategy is to estimate disparities after adjusting for observed covariates, typically with a regression model. This approach, however, suffers from two key statistical challenges. First, omitted-variable bias can skew results if the model does not adjust for all relevant factors; second, and conversely, included-variable bias -- a lesser-known phenomenon -- can skew results if the set of covariates includes irrelevant factors. Here we introduce a new, three-step statistical method, which we call risk-adjusted regression, to address both concerns in settings where decision makers have clearly measurable objectives. In the first step, we use all available covariates to estimate the value, or inversely, the risk, of taking a certain action, such as approving a loan application or hiring a job candidate. Second, we measure disparities in decisions after adjusting for these risk estimates alone, mitigating the problem of included-variable bias. Finally, in the third step, we assess the sensitivity of results to potential mismeasurement of risk, addressing concerns about omitted-variable bias. To do so, we develop a novel, non-parametric sensitivity analysis that yields tight bounds on the true disparity in terms of the average gap between true and estimated risk -- a single interpretable parameter that facilitates credible estimates. We demonstrate this approach on a detailed dataset of 2.2 million police stops of pedestrians in New York City, and show that traditional statistical tests of discrimination can substantially underestimate the magnitude of disparities.

研究动机与目标

  • 解决差异影响估计中的遗漏变量偏差,其中与群体成员身份相关的未测量混杂因素会扭曲结果。
  • 纠正包含无关或代理变量(如邮政编码、声线)导致的包含变量偏差,这些变量与种族相关,从而扭曲歧视估计。
  • 开发一种方法,仅通过调整估计风险来隔离真实差异,将受保护属性从风险建模中排除。
  • 提供一种非参数敏感性分析,基于真实风险与估计风险之间的平均绝对差值(ε)来界定真实差异的边界。
  • 证明传统回归方法在现实决策系统中可能显著低估差异影响的实际程度。

提出的方法

  • 步骤1:使用除受保护属性(如种族、性别)外的所有可用协变量,通过预测模型估计结果(如携带武器)的先验风险。
  • 步骤2:通过仅将决策结果(如搜身)对估计风险进行回归,来估计差异,从而消除无关协变量带来的包含变量偏差。
  • 步骤3:进行非参数敏感性分析,基于真实风险与估计风险之间的平均绝对差值(ε)来界定真实差异的边界,ε为单一可解释参数。
  • 敏感性分析通过在群体层面总风险上进行网格搜索,并使用精度参数(γ = 0.01个百分点)的优化过程,计算出差异影响的紧致边界。
  • 该方法采用自举法置信区间(N=1000)来评估不确定性,模型校准通过AUC(81%)和命中率比较进行验证。
  • rar R 包在 CRAN 上实现了敏感性分析,所有数据和代码均可公开获取以供复现。

实验结果

研究问题

  • RQ1在现实决策系统中,由于遗漏变量和包含变量偏差,传统回归模型在多大程度上低估了差异影响?
  • RQ2当受保护属性被排除在风险建模之外时,风险调整回归在多大程度上能减少差异影响估计中的偏差?
  • RQ3风险估计误差对估计差异影响程度有何影响?该影响能否被定量界定?
  • RQ4当仅根据武器携带的估计风险进行调整,而不将种族纳入风险模型时,不同种族群体之间的搜身率差异如何变化?
  • RQ5在合理范围内的风险估计误差水平下,风险调整回归的结果在多大程度上保持稳健?

主要发现

  • 风险调整回归方法揭示,即使在估计携带武器风险相当的情况下,非裔和西班牙裔行人被搜身的概率仍显著高于白人行人,表明存在显著的差异影响。
  • 包含所有搜身前协变量的传统“全量回归”模型因包含无关变量而产生偏差,导致真实差异被低估。
  • 该模型在纽约市警察局数据的后半部分实现了81%的AUC,表明其对武器携带的预测性能较强。
  • 模型预测的命中率在不同人口群体和警区中与实际命中率高度一致,证实了风险模型的良好校准性。
  • 即使在严重混杂(ε = 0.7个百分点)的情况下,估计的差异影响依然显著,有44%的差异仍无法由风险估计解释。
  • 敏感性分析表明,差异在广泛合理的风险估计误差水平下依然持续存在,且95%置信区间紧密收敛。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。