[论文解读] Causal Inference Under Network Interference: A Framework for Experiments on Social Networks
本文提出了一种贝叶斯框架,用于在存在干扰的社会网络中进行因果推断,其中个体的结果依赖于其邻居的处理情况。该框架使用网络潜在结果和合理的插补方法来估计主要效应、同伴效应和总效应,表明在模型误设情况下,平衡的实验设计和网络协变量的引入可提高估计精度。
No man is an island, as individuals interact and influence one another daily in our society. When social influence takes place in experiments on a population of interconnected individuals, the treatment on a unit may affect the outcomes of other units, a phenomenon known as interference. This thesis develops a causal framework and inference methodology for experiments where interference takes place on a network of influence (i.e. network interference). In this framework, the network potential outcomes serve as the key quantity and flexible building blocks for causal estimands that represent a variety of primary, peer, and total treatment effects. These causal estimands are estimated via principled Bayesian imputation of missing outcomes. The theory on the unconfoundedness assumptions leading to simplified imputation highlights the importance of including relevant network covariates in the potential outcome model. Additionally, experimental designs that result in balanced covariates and sizes across treatment exposure groups further improve the causal estimate, especially by mitigating potential outcome model mis-specification. The true potential outcome model is not typically known in real-world experiments, so the best practice is to account for interference and confounding network covariates through both balanced designs and model-based imputation. A full factorial simulated experiment is formulated to demonstrate this principle by comparing performance across different randomization schemes during the design phase and estimators during the analysis phase, under varying network topology and true potential outcome models. Overall, this thesis asserts that interference is not just a nuisance for analysis but rather an opportunity for quantifying and leveraging peer effects in real-world experiments.
研究动机与目标
- 解决社会网络实验中干扰带来的挑战,即一个单位的处理会影响其他单位。
- 开发一种灵活的因果框架,将结果建模为个体处理和同伴处理的函数。
- 通过将网络结构和协变量整合到潜在结果模型中,改进因果估计。
- 证明干扰不仅是干扰因素,更是量化和利用同伴效应的机会。
- 提供设计与分析指南,通过平衡随机化和协变量调整,增强对模型误设的鲁棒性。
提出的方法
- 在具有网络干扰的稳定单位处理值假设(SUTNVA)下定义网络潜在结果,允许存在同伴效应。
- 引入主要效应、k个受处理邻居、固定分配和影响网络操控效应的因果估计算量。
- 应用贝叶斯插补方法,利用包含网络协变量的模型来估计缺失的潜在结果。
- 在存在网络干扰的情况下推导出无偏性假设,表明对相关网络结构进行条件化可实现有效的推断。
- 使用混合成员关系块模型生成具有社区结构、混合成员关系和稀疏性的现实网络拓扑。
- 采用MCMC和MCEM进行后验推断,性能边界基于成员关系参数的费舍尔信息推导得出。
实验结果
研究问题
- RQ1在存在网络干扰的情况下,如何正式定义因果估计算量?
- RQ2在存在干扰时,确保有效因果推断的条件是什么?如何实现无偏性?
- RQ3实验设计选择(如随机化方案和协变量平衡)在模型误设下如何影响估计精度?
- RQ4网络拓扑结构(如稀疏性、社区结构)对因果估计器性能有何影响?
- RQ5在存在干扰的场景下,基于模型的插补结合平衡设计是否能优于标准方法?
主要发现
- 在潜在结果模型中引入网络协变量可显著提高估计精度,因为满足了无偏性假设。
- 通过在处理暴露组之间平衡协变量的重随机化方法可减少偏差并提升性能,尤其在模型误设时效果更明显。
- 混合成员关系块模型(HMMB)能有效捕捉重叠社区和异质连接等现实网络特征。
- 随着稀疏性增加和跨社区交互增强,性能下降,但当模型设定正确时,贝叶斯插补仍保持稳健。
- 模拟实验中后验区间宽度和覆盖概率表明,当设计与建模策略精心对齐时,估计器表现最佳。
- 全因子模拟研究识别出,重随机化与基于模型的插补相结合,在不同网络拓扑和真实结果模型下均能产生最鲁棒的因果估计。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。