[论文解读] Regression adjustments for estimating the global treatment effect in experiments with interference
本文提出了一种回归调整估计量,用于在存在干扰的随机化实验中无偏估计全局平均处理效应(GATE),使用处理分配函数作为协变量,并通过机器学习实现灵活建模。该方法在保持潜在结果可忽略性假设下,相较于现有逆倾向得分加权方法,提升了偏差校正效果,并通过自助抽样实现有效的推断。
Standard estimators of the global average treatment effect can be biased in the presence of interference. This paper proposes regression adjustment estimators for removing bias due to interference in Bernoulli randomized experiments. We use a fitted model to predict the counterfactual outcomes of global control and global treatment. Our work differs from standard regression adjustments in that the adjustment variables are constructed from functions of the treatment assignment vector, and that we allow the researcher to use a collection of any functions correlated with the response, turning the problem of detecting interference into a feature engineering problem. We characterize the distribution of the proposed estimator in a linear model setting and connect the results to the standard theory of regression adjustments under SUTVA. We then propose an estimator that allows for flexible machine learning estimators to be used for fitting a nonlinear interference functional form. We propose conducting statistical inference via bootstrap and resampling methods, which allow us to sidestep the complicated dependences implied by interference and instead rely on empirical covariance structures. Such variance estimation relies on an exogeneity assumption akin to the standard unconfoundedness assumption invoked in observational studies. In simulation experiments, our methods are better at debiasing estimates than existing inverse propensity weighted estimators based on neighborhood exposure modeling. We use our method to reanalyze an experiment concerning weather insurance adoption conducted on a collection of villages in rural China.
研究动机与目标
- 解决由于随机化实验中存在干扰而导致的全局平均处理效应(GATE)估计偏差问题。
- 通过从处理分配函数构造协变量,将通常在SUTVA下使用的回归调整技术扩展到存在干扰的情境。
- 利用机器学习估计器对均值函数进行灵活的非线性建模,以捕捉干扰效应。
- 在可忽略性假设下发展有效的统计推断,避免依赖复杂的解析方差公式。
- 通过模拟实验和真实数据表明,与现有逆倾向得分加权方法相比,该方法在偏差和精度方面均有改进。
提出的方法
- 将回归调整的协变量构造为处理分配向量的函数,通过与响应变量相关的用户定义函数捕捉干扰模式。
- 采用线性模型框架刻画所提估计量的分布,并将其与SUTVA下的标准回归调整理论相联系。
- 提出一种灵活的估计量,利用机器学习模型拟合非线性干扰函数形式,从而在复杂情境中提升模型拟合度。
- 采用自助抽样和重抽样方法进行方差估计,利用经验协方差结构以处理由干扰引起的依赖性。
- 依赖于类似于观察性研究中无偏性的可忽略性假设,确保在未观测混杂因素不存在时可实现有效推断。
- 将该方法应用于中国农村天气保险实验的真实数据,展示了其实际应用价值。
实验结果
研究问题
- RQ1当干扰违反SUTVA时,回归调整能否有效降低GATE估计的偏差?
- RQ2如何将处理分配函数用作协变量,以一种可推广至标准回归调整的方式建模干扰?
- RQ3与解析方差公式相比,基于自助抽样的GATE估计量推断在干扰条件下的表现如何?
- RQ4与基于暴露模型的逆倾向得分加权方法相比,所提方法在偏差和精度方面表现如何?
- RQ5在存在复杂非线性干扰模式的情境下,机器学习模型能否提升估计精度?
主要发现
- 在模拟实验中,所提出的回归调整估计量相较于逆倾向得分加权估计量显著降低了偏差。
- 在对一项中国农村天气保险采纳实验的再分析中,线性回归调整和逻辑回归调整的GATE估计值分别为0.1218和0.1197,标准误分别为0.0561和0.0559。
- 标准误估计值较宽,提示在解释处理效应时应保持谨慎;同时发现Hájek估计量的保守方差估计值不稳定,且常大于1。
- 该方法通过自助抽样实现有效推断,避免了在干扰条件下对复杂解析方差计算的依赖。
- 该方法将干扰检测视为特征工程问题,允许研究者通过函数选择融入领域知识。
- 使用聚类自助或稳健标准误的残差诊断方法,有助于检测均值模型中未捕捉到的剩余干扰。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。