[论文解读] Combining Observational and Experimental Datasets Using Shrinkage Estimators
本文提出了收缩估计量,通过结合实验数据与观察数据,降低在未测量混杂因素导致观察估计偏差时因果推断的风险。通过向实验估计量进行 Stein 型收缩,作者提出了两种新估计量——$κ_1$ 和 $κ_2$——在温和条件下,即使观察数据因未测量混杂因素而存在偏差,其有限样本风险也低于仅使用实验数据的估计量。
We consider the problem of combining data from observational and experimental sources to make causal conclusions. This problem is increasingly relevant, as the modern era has yielded passive collection of massive observational datasets in areas such as e-commerce and electronic health. These data may be used to supplement experimental data, which is frequently expensive to obtain. In Rosenman et al. (2018), we considered this problem under the assumption that all confounders were measured. Here, we relax the assumption of unconfoundedness. To derive combined estimators with desirable properties, we make use of results from the Stein Shrinkage literature. Our contributions are threefold. First, we propose a generic procedure for deriving shrinkage estimators in this setting, making use of a generalized unbiased risk estimate. Second, we develop two new estimators, prove finite sample conditions under which they have lower risk than an estimator using only experimental data, and show that each achieves a notion of asymptotic optimality. Third, we draw connections between our approach and results in sensitivity analysis, including proposing a method for evaluating the feasibility of our estimators.
研究动机与目标
- 通过结合观察数据(高偏差、低方差)与实验数据(低偏差、高方差),解决因果推断中的偏差-方差权衡问题。
- 建立一个通用框架,用于推导平衡有偏观察估计量与无偏实验估计量的收缩估计量。
- 在存在未测量混杂因素的情况下,建立有限样本条件下所提估计量风险低于仅使用实验数据估计量的条件。
- 将收缩方法与敏感性分析相联系,实现对未测量混杂因素影响的稳健性评估。
- 提供实用的实施指南,包括方差估计与倾向得分调整。
提出的方法
- 将 Strawderman(1975)的无偏风险估计方法推广至异方差、多变量设定,推导组合估计量的风险估计。
- 提出两种收缩估计量:$κ_1$ 在各分层中应用统一的收缩因子,$κ_2$ 使用基于方差加权的收缩因子。
- 通过最小化广义无偏风险估计,推导出估计量的函数形式,确保最优收缩至实验估计量。
- 在考虑数据双重使用(用于收缩与估计)的前提下,采用偏差校正优化程序,实现收缩因子的数据驱动估计。
- 采用 Zhao 等人提出的敏感性分析框架,估计隐含的 $Γ$ 参数,以评估对未测量混杂因素的稳健性。
- 在不同协变量分布与偏差水平的模拟中应用估计量,与 Green 和 Strawderman 的估计量($δ_1$,$δ_2$)进行性能比较。
实验结果
研究问题
- RQ1在何种条件下,通过收缩方法结合观察与实验数据,可使风险低于仅使用实验数据?
- RQ2当观察数据受未测量变量混杂影响时,如何构建收缩估计量以最优平衡偏差与方差?
- RQ3在协变量分布不同的设定下,所提估计量的有限样本表现与现有收缩方法相比如何?
- RQ4如何将敏感性分析与收缩估计结合,以评估对未测量混杂因素的稳健性?
- RQ5在可靠实施所提估计量时,需考虑哪些实际问题,如方差估计与倾向得分调整?
主要发现
- 在有限样本条件下,所提估计量 $κ_1$ 与 $κ_2$ 的风险低于仅使用实验数据的估计量,尤其当观察数据具有中等偏差与高精度时表现更优。
- 在协变量分布相同的模拟中,$κ_1$ 的风险始终低于 Green 与 Strawderman 的 $δ_1$,尤其当分层数 $K \geq 4$ 时更为显著。
- 当观察与实验数据的协变量分布不同时,$κ_1$ 仍具有效性,但 $δ_1$ 稍微更具稳健性,表明风险降低与分布敏感性之间存在权衡。
- 随着分层数 $K \to \infty$,估计量 $κ_1$ 与 $κ_2$ 达到渐近最优性,展现出长期效率。
- 与敏感性分析的关联使研究者能够估计隐含的 $Γ$ 参数,从而提供对未测量混杂因素影响的定量稳健性度量。
- 该方法具有可扩展性,可进一步整合辅助信息或阈值规则,以在高偏差分层中减少收缩程度。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。