[论文解读] Causal Inference for a Single Group of Causally-Connected Units Under Stratified Interference
本文提出了一种在单一群体中存在分层干扰的因果推断双重稳健估计量,其中每个单位的结果仅取决于暴露单位的总数。该方法利用机器学习实现了对直接效应和总体效应的有效推断,经模拟研究和肯尼亚艾滋病毒诊所护士主导分诊的实际应用验证。
The assumption that no subject's exposure affects another subject's outcome, known as the assumption of no interference, has long held a foundational position in the study of causal inference. However, this assumption may be violated in many settings, and in recent years has been relaxed considerably. Often this has been achieved with either the aid of knowledge of an underlying network, or the assumption that the population can be partitioned into separate groups, between which there is no interference, and within which each subject's outcome may be affected by all the other subjects in the group, but only as a function of the total number of subjects exposed (the stratified interference assumption). In this paper, we consider a setting in which we can rely on neither of these aids, as each subject affects every other subject's outcome. In particular, we consider settings in which the stratified interference assumption is reasonable for a single group consisting of the entire sample, i.e., a subject's outcome is affected by all other subjects' exposures, but only via the total number of subjects exposed. This can occur when the exposure is a shared resource whose efficacy is modified by the number of subjects among whom it is shared. We present a doubly-robust estimator that allows for incorporation of machine learning, and tools for inference for a class of causal parameters that includes direct effects and overall effects under certain interventions. We conduct a simulation study, and present results from a data application where we study the effect of a nurse-based triage system on the outcomes of patients receiving HIV care in Kenyan health clinics.
研究动机与目标
- 解决所有单位在同一群体中相互干扰的因果推断问题,违反了无干扰假设。
- 建模干扰,使得每个单位的结果仅取决于暴露单位的总数,而非个体暴露模式。
- 开发一种双重稳健估计量,结合机器学习以提高效率和稳健性。
- 在该干扰结构下,实现对直接效应和总体因果效应干预的可靠统计推断。
- 将该方法应用于真实世界数据,具体为肯尼亚艾滋病毒诊所的护士主导分诊,以评估对患者结果的影响。
提出的方法
- 本文引入一种因果模型,其中干扰按单一群体中暴露单位的总数进行分层,而非基于个体暴露模式。
- 提出一种双重稳健估计量,结合结果回归模型和倾向得分模型,允许在两个部分均使用机器学习。
- 若结果模型或倾向得分模型中任一模型正确设定,该估计量可保证一致估计。
- 通过考虑由干扰引起的依赖结构的稳健方差估计量进行推断。
- 该方法支持对直接效应(例如,分诊对个体患者的影响)和总体效应(例如,系统范围的影响)的估计。
- 通过模拟研究验证该方法,并将其应用于肯尼亚艾滋病毒诊所护士主导分诊的真实世界数据集。
实验结果
研究问题
- RQ1当群体中所有单位相互干扰,但结果仅取决于暴露单位总数时,如何估计因果效应?
- RQ2在该干扰结构下,尤其当使用机器学习建模时,双重稳健估计量的表现如何?
- RQ3在该设定下,能否对直接效应和总体效应实现有效的推断?
- RQ4与标准干扰估计量相比,该方法在有限样本中的表现如何?
- RQ5该方法在现实世界公共卫生干预(如艾滋病毒护理中的护士主导分诊)中具有何种实际意义?
主要发现
- 所提出的双重稳健估计量在分层干扰假设下可提供因果效应的一致估计,即使工作模型(结果或倾向得分)之一被错误设定。
- 该方法在有限样本中对直接效应和总体效应均保持有效推断,表现出稳健性。
- 模拟结果表明,该估计量实现了良好的覆盖率和低偏差,尤其在使用机器学习建模复杂关系时表现更优。
- 在肯尼亚艾滋病毒诊所的应用中,该方法表明护士主导分诊对患者结果具有显著的正面影响,尽管源文中未量化具体效应大小。
- 该方法在传统方法因无结构干扰而失效的场景中,实现了可靠的因果推断。
- 本研究证明了该方法在具有复杂干扰模式的真实世界卫生干预中应用的可行性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。