Skip to main content
QUICK REVIEW

[论文解读] Strategic Instrumental Variable Regression: Recovering Causal Relationships From Strategic Responses

Keegan Harris, Dung Daniel T. Ngo|arXiv (Cornell University)|Jul 12, 2021
Experimental Behavioral Economics Studies参考文献 40被引用 5
一句话总结

本文提出了一种新颖的方法,通过利用模型更新所引发的战略性响应作为工具变量,来恢复机器学习中的因果关系。通过将部署的模型视为影响可观测特征但不直接影响结果的工具变量,该方法即使在存在未观测混杂因素的情况下也能实现因果推断,从而提升决策系统中的公平性、结果质量及预测风险。

ABSTRACT

In settings where Machine Learning (ML) algorithms automate or inform consequential decisions about people, individual decision subjects are often incentivized to strategically modify their observable attributes to receive more favorable predictions. As a result, the distribution the assessment rule is trained on may differ from the one it operates on in deployment. While such distribution shifts, in general, can hinder accurate predictions, our work identifies a unique opportunity associated with shifts due to strategic responses: We show that we can use strategic responses effectively to recover causal relationships between the observable features and outcomes we wish to predict, even under the presence of unobserved confounding variables. Specifically, our work establishes a novel connection between strategic responses to ML models and instrumental variable (IV) regression by observing that the sequence of deployed models can be viewed as an instrument that affects agents' observable features but does not directly influence their outcomes. We show that our causal recovery method can be utilized to improve decision-making across several important criteria: individual fairness, agent outcomes, and predictive risk. In particular, we show that if decision subjects differ in their ability to modify non-causal attributes, any decision rule deviating from the causal coefficients can lead to (potentially unbounded) individual-level unfairness.

研究动机与目标

  • 解决个体在预测模型影响下战略性地修改其可观测特征时,机器学习中因果推断的挑战。
  • 纠正高风险决策系统中特征与结果之间因果关系因未观测混杂因素而产生的扭曲。
  • 开发一种方法,利用部署模型的序列作为有效工具变量,以恢复真实的因果系数。
  • 通过确保决策规则与因果关系保持一致,改善个体公平性、代理结果及预测风险。
  • 证明当不同群体的战略响应能力不同时,偏离因果系数可能导致个体层面的无界不公平性。

提出的方法

  • 该方法将战略性响应建模为部署评估规则序列的函数,将这些规则视为工具变量。
  • 假设模型序列影响代理的可观测特征,但不直接影响其结果,满足工具变量回归中的排除限制条件。
  • 通过两阶段最小二乘法(2SLS)回归恢复因果系数,其中模型序列为工具变量,可观测特征为内生回归变量。
  • 通过将未观测混杂因素建模为同时影响特征和结果的潜变量,而工具变量与这些混杂因素保持独立,从而处理未观测混杂因素。
  • 该框架在简化的大学录取场景中得到验证,并扩展至信用评分和事故风险预测,展示了其在不同领域中的鲁棒性。
  • 理论分析表明,若使用非因果决策规则,当个体在改变非因果特征的能力上不同时,可能导致无界的不公平性。

实验结果

研究问题

  • RQ1是否可以利用机器学习模型引发的战略性响应来恢复可观测特征与结果之间的因果关系?
  • RQ2在存在未观测混杂因素的情况下,部署模型的序列在何种条件下可作为有效的工具变量?
  • RQ3通过战略性响应实现的因果恢复如何改善决策系统中的公平性、代理结果及预测风险?
  • RQ4当个体在修改非因果特征的能力上不同时,使用非因果决策规则会带来何种后果?
  • RQ5该方法能否推广至现实世界应用,如大学录取和信用评分?

主要发现

  • 部署模型的序列可作为有效工具变量,即使未观测混杂因素同时影响特征和结果,也能实现因果恢复。
  • 使用非因果决策规则会导致个体层面的不公平性无界增长,当代理在战略响应能力上存在差异时尤为明显。
  • 所提出的方法通过使决策规则与真实因果关系对齐,而非与虚假相关性对齐,从而改善了个体公平性。
  • 在1D模拟中,该方法即使在存在混杂和战略性操纵的数据上进行训练,也能成功恢复真实的因果系数 θ* = 1。
  • 在由战略性行为引发的分布偏移下,该方法在保持鲁棒性方面优于标准预测模型。
  • 理论分析确认,模型更新中使用的梯度必须考虑特征对模型的内生依赖关系,否则可能导致错误的优化方向。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。