Skip to main content
QUICK REVIEW

[论文解读] Assessing lack of common support in causal inference using Bayesian nonparametrics: Implications for evaluating the effect of breastfeeding on children's cognitive outcomes

Jennifer Hill, Yu‐Sung Su|arXiv (Cornell University)|Nov 28, 2013
Advanced Causal Inference Techniques被引用 4
一句话总结

本文提出一种基于贝叶斯非参数方法的BART(贝叶斯加性回归树)方法,用于在因果推断中识别缺乏共同因果支持的单位,通过利用反事实结果的后验不确定性来检测不可靠的推断。结果表明,与传统倾向得分方法相比,基于BART的剔除规则在高维设置下能更好地保留有效单位,并更准确地反映真实的协变量-结果关系,尤其在存在复杂交互作用的情况下表现更优。

ABSTRACT

Causal inference in observational studies typically requires making comparisons between groups that are dissimilar. For instance, researchers investigating the role of a prolonged duration of breastfeeding on child outcomes may be forced to make comparisons between women with substantially different characteristics on average. In the extreme there may exist neighborhoods of the covariate space where there are not sufficient numbers of both groups of women (those who breastfed for prolonged periods and those who did not) to make inferences about those women. This is referred to as lack of common support. Problems can arise when we try to estimate causal effects for units that lack common support, thus we may want to avoid inference for such units. If ignorability is satisfied with respect to a set of potential confounders, then identifying whether, or for which units, the common support assumption holds is an empirical question. However, in the high-dimensional covariate space often required to satisfy ignorability such identification may not be trivial. Existing methods used to address this problem often require reliance on parametric assumptions and most, if not all, ignore the information embedded in the response variable. We distinguish between the concepts of "common support" and common causal support." We propose a new approach for identifying common causal support that addresses some of the shortcomings of existing methods. We motivate and illustrate the approach using data from the National Longitudinal Survey of Youth to estimate the effect of breastfeeding at least nine months on reading and math achievement scores at age five or six. We also evaluate the comparative performance of this method in hypothetical examples and simulations where the true treatment effect is known.

研究动机与目标

  • 解决在高维协变量空间中估计因果效应时,识别缺乏共同因果支持单位的挑战。
  • 开发一种利用BART后验分布中的结果信息来指导剔除决策的方法,而非仅依赖处理分配模型。
  • 提供一种更稳定且基于实证的方法,用于在剔除重叠不足的单位后定义推断样本。
  • 比较基于BART的剔除规则与标准倾向得分方法在单位保留率和估计准确性方面的表现。
  • 通过真实数据示例(母乳喂养持续时间与儿童认知结果)说明该方法,突出其在推断中的实际差异。

提出的方法

  • 使用贝叶斯加性回归树(BART)建模潜在结果,为每个单位生成事实结果和反事实结果的个体后验分布。
  • 通过比较反事实结果估计的精度(后验标准差)与观测结果估计的精度,识别缺乏共同因果支持的单位。
  • 应用一种规则:当反事实结果的后验不确定性超过阈值(例如,高于平均不确定性1个标准差)时,将该单位标记为剔除对象。
  • 使用回归树对被剔除单位进行特征分析,以理解导致剔除决策的协变量模式。
  • 在模拟和真实数据中,将基于BART的剔除规则与标准倾向得分方法(如卡钳匹配、倾向得分五等分法)进行比较。
  • 不仅将BART后验分布用于因果效应估计,还将其作为诊断工具,评估单个单位推断的可靠性。

实验结果

研究问题

  • RQ1在不依赖处理分配的参数模型前提下,如何可靠地识别高维协变量空间中缺乏共同因果支持的单位?
  • RQ2在共同支持评估中引入结果信息,能在多大程度上提升因果推断的有效性和稳定性?
  • RQ3基于BART的剔除规则与标准倾向得分方法相比,在单位保留率和估计准确性方面表现如何?
  • RQ4在母乳喂养与儿童认知关系的真实世界示例中,使用BART方法与倾向得分方法在估计治疗效应上存在哪些实质性差异?
  • RQ5如何有意义地描述因缺乏共同支持而被剔除单位的特征?

主要发现

  • 在母乳喂养的真实数据示例中,基于BART的剔除规则保留了所有单位,而倾向得分方法剔除了大量单位,尤其集中在低教育水平、低AFQT得分的母亲中。
  • BART方法识别出更少的缺乏共同因果支持的单位,表明其比标准倾向得分方法更不保守,且更能反映真实数据结构。
  • 在已知真实治疗效应的模拟中,BART的1个标准差剔除规则表现更优,既最小化了不必要的剔除,又保持了估计准确性。
  • 该方法揭示,某些在倾向得分模型中曾被视为混杂因素的变量,对结果的预测能力很弱,凸显了基于结果的信息评估支持的重要性。
  • 对被剔除单位的回归树分析表明,基于BART的剔除决策由有意义的协变量模式驱动,而倾向得分方法的剔除往往缺乏明确的实质性依据。
  • 研究发现,仅依赖倾向得分会导致估计不稳定和结果不一致,尤其在匹配分析中表现明显,而基于BART的推断则产生了更一致且更具实质性合理性的效应。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。