[论文解读] A systematic investigation of classical causal inference strategies under mis-specification due to network interference
本文系统研究了网络干扰下的因果推断,提出了一种基于暴露邻域的框架来建模干扰效应。研究发现,由于未知的干扰模式,经典估计量(如均值差)可能产生严重偏差;而霍尔维茨-汤普森估计量仅在已知倾向得分时才无偏,但在许多设计下仍存在非可接受性,因此提出了改进的模型辅助和设计驱动策略,以更准确估计平均和总处理效应。
We systematically investigate issues due to mis-specification that arise in estimating causal effects when (treatment) interference is informed by a network available pre-intervention, i.e., in situations where the outcome of a unit may depend on the treatment assigned to other units. We develop theory for several forms of interference through the concept of exposure neighborhood, and develop the corresponding semi-parametric representation for potential outcomes as a function of the exposure neighborhood. Using this representation, we extend the definition of two popular classes of causal estimands, marginal and average causal effects, to the case of network interference. We characterize the bias and variance one incurs when combining classical randomization strategies (namely, Bernoulli, Completely Randomized, and Cluster Randomized designs) and estimators (namely, difference-in-means and Horvitz-Thompson) used to estimate average treatment effect and on the total treatment effect, under misspecification due to interference. We illustrate how difference-in-means estimators can have arbitrarily large bias when estimating average causal effects, depending on the form and strength of interference, which is unknown at design stage. Horvitz-Thompson (HT) estimators are unbiased when the correct weights are specified. Here, we derive the HT weights for unbiased estimation of different estimands, and illustrate how they depend on the design, the form of interference, which is unknown at design stage, and the estimand. More importantly, we show that HT estimators are in-admissible for a large class of randomization strategies, in the presence of interference. We develop new model-assisted and model-dependent strategies to improve HT estimators, and we develop new randomization strategies for estimating the average treatment effect and total treatment effect.
研究动机与目标
- 解决由于预存网络结构导致的处理干扰所带来的因果推断挑战。
- 利用暴露邻域模型,正式定义在网络干扰下因果效应的估计量——特别是边际效应和平均因果效应。
- 在干扰误设条件下,刻画经典随机化设计(伯努利、完全随机设计、聚类设计)和估计量(均值差、霍尔维茨-汤普森)的偏差与方差。
- 证明在常见干扰场景下,霍尔维茨-汤普森估计量存在非可接受性,并提出改进的模型辅助与设计驱动替代方案。
- 开发新型随机化策略(如独立集设计与聚类随机化设计),以在干扰存在时估计直接效应与总处理效应。
提出的方法
- 引入‘暴露邻域’概念,以建模单位潜在结果如何依赖于其网络邻居的处理分配。
- 建立潜在结果的半参数表示,作为暴露邻域的函数,从而实现对直接效应、干扰效应与交互效应的分解。
- 通过结构模型与暴露模型,将边际效应与平均因果效应估计量扩展至网络干扰情境。
- 推导霍尔维茨-汤普森权重以实现无偏估计,表明其依赖于设计、干扰结构与估计量,需事先掌握未知的倾向得分。
- 提出广义线性估计量、模型相关估计量与模型辅助估计量,以在模型误设下改进霍尔维茨-汤普森估计量的表现。
- 引入新型随机化设计:独立集设计(用于直接效应估计)与聚类随机化设计(用于总处理效应估计),均针对减少干扰引起的偏差进行优化。
实验结果
研究问题
- RQ1网络干扰如何使经典因果推断策略(如均值差估计量)失效,会产生何种形式的偏差?
- RQ2在干扰存在时,霍尔维茨-汤普森估计量在何种条件下对平均与总处理效应是无偏的,又在何种情况下表现出非可接受性?
- RQ3当干扰存在但未知时,不同随机化设计(伯努利、完全随机设计、聚类设计、独立集设计)在偏差与方差方面的表现如何?
- RQ4在干扰误设下,模型辅助或模型相关估计量是否能在均方误差方面优于霍尔维茨-汤普森估计量?
- RQ5能否设计出新型随机化策略,以在干扰存在且未知时,改善直接效应与总处理效应的估计?
主要发现
- 当干扰结构与暴露模型未知时,均值差估计量在估计平均因果效应时可能表现出任意大的偏差,尤其在干扰模式复杂时更为显著。
- 霍尔维茨-汤普森估计量仅在已知正确倾向得分时才无偏,但在许多标准设计下因方差过高且对模型误设敏感而表现出非可接受性。
- 本文推导出显式的霍尔维茨-汤普森权重,其依赖于干扰图、暴露模型与估计量,但通常在缺乏网络结构先验知识时难以计算。
- 独立集设计通过将自我单位与干扰隔离,可确保直接效应的无偏估计,但倾向得分需借助蒙特卡洛方法计算,因选择概率复杂。
- 聚类随机化设计通过最大化相关潜在结果数量,有效估计总处理效应,但若不暴露干扰性潜在结果,则无法估计直接效应。
- 模型辅助与模型相关估计量在均方误差方面显著优于霍尔维茨-汤普森估计量,尤其当真实干扰结构近似线性或可加时。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。