Skip to main content
QUICK REVIEW

[论文解读] A causal inference framework for spatial confounding

Brian Gilbert, Abhirup Datta|arXiv (Cornell University)|Dec 30, 2021
Spatial and Panel Data Analysis被引用 10
一句话总结

该论文提出了一种模型无关的因果推断框架,用于处理空间混杂因素,将其定义为一种未观测到的空间结构化混杂因子。该框架引入了两个关键识别假设——混杂因子可作为空间坐标的可测函数,且暴露变量中存在非空间变异——并开发了一种双重机器学习(DML)估计器,确保在空间依赖条件下具有稳健性且渐近正态,模拟和一项关于PM2.5与出生体重的真实世界研究均表明其优于传统方法。

ABSTRACT

Over the past few decades, addressing "spatial confounding" has become a major topic in spatial statistics. However, the literature has provided conflicting definitions, and many proposed solutions are tied to specific analysis models and do not address the issue of confounding as it is understood in causal inference. We offer an analysis-model-agnostic definition of spatial confounding as the existence of an unmeasured causal confounder variable with a spatial structure. We present a causal inference framework for nonparametric identification of the causal effect of a continuous exposure on an outcome in the presence of spatial confounding. In particular, we identify two critical additional assumptions that allow the use of the spatial coordinates as a proxy for the unmeasured spatial confounder: the measurability of the confounder as a function of space, which is required for conditional ignorability to hold, and the presence of a non-spatial component in the exposure, required for positivity to hold. We also propose studying a causal estimand based on a "shift intervention" that requires less stringent identifying assumptions than traditional estimands. We then turn to estimation and focus on "double machine learning" (DML), a procedure in which flexible models are used to regress both the exposure and outcome variables on confounders to arrive at a causal estimator with favorable robustness properties and convergence rates. This procedure avoids restrictive assumptions, such as linearity and effect homogeneity, which are typically made in spatial models and which can lead to bias when violated. We demonstrate the advantages of the DML approach analytically and via extensive simulation studies. We apply our methods and reasoning to a study of the effect of fine particulate matter exposure during pregnancy on birthweight in California.

研究动机与目标

  • 以模型无关的方式定义空间混杂,即存在一种具有空间结构的未观测混杂因子。
  • 识别在存在空间混杂时实现有效因果识别所需的最小非参数假设。
  • 提出一种偏移估计量,其对正性假设的要求弱于标准剂量-反应估计量。
  • 提出并证明一种双重机器学习(DML)方法,以在空间依赖条件下实现灵活、稳健的估计。
  • 通过模拟研究和一项关于孕期PM2.5暴露与出生体重的真实世界应用,展示该方法优于现有方法。

提出的方法

  • 将空间混杂定义为存在一种与分析模型无关的未观测混杂因子,且具有空间结构。
  • 引入两个关键识别假设:(1) 混杂因子是空间坐标的可测函数,从而实现条件忽略性;(2) 暴露变量中存在非空间变异,从而确保正性。
  • 提出一种偏移估计量,作为完整剂量-反应曲线的更稳健替代,其对正性假设的要求更弱。
  • 通过使用灵活模型将结果和暴露变量对混杂因子进行回归,然后通过残差化估计因果效应,应用双重机器学习(DML)。
  • 在空间依赖条件下证明了DML估计器的渐近正态性,确保了有效的统计推断。
  • 通过模拟研究和对加利福尼亚州PM2.5暴露与出生体重的真实世界分析,验证了该方法的有效性。

实验结果

研究问题

  • RQ1在存在空间混杂时,哪些非参数假设可使因果效应被识别?
  • RQ2如何在不依赖严格参数模型的前提下,将空间坐标用作未观测空间混杂因子的代理变量?
  • RQ3在使用空间坐标作为混杂因子代理变量时,哪些条件可确保正性?
  • RQ4与标准估计量相比,所提出的偏移估计量在识别假设和可解释性方面有何差异?
  • RQ5双重机器学习是否能在具有复杂依赖结构的空间数据中提供稳健、高效且渐近正态的推断?

主要发现

  • 基于BART和样条模型的DML估计器在高噪声模拟中实现了最低的均方误差(5.621×10⁻⁵)和最高的覆盖区间(84%)。
  • 即使在非线性和异质效应条件下,DML方法也显著降低了偏差和均方误差,优于单重机器学习方法。
  • 偏移估计量对正性假设的要求弱于完整剂量-反应估计,从而在实践中提升了稳健性。
  • 该方法在非线性关系和效应异质性较强的场景下,优于传统空间模型(如gSEM、spatial+和RSR),表现更优。
  • 在真实世界应用中,DML框架对孕期PM2.5暴露与出生体重因果效应的推断比传统方法更可靠。
  • 对平滑混杂表面的轻微偏离对性能影响极小,表明该方法对平滑性假设的轻微违反具有鲁棒性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。