[论文解读] Minimax Rates and Adaptivity in Combining Experimental and Observational Data
本文建立了结合随机对照试验(RCT)和观察性数据以估计平均处理效应的极小化极大最优速率,提出了一种完全自适应的锚定阈值估计器,可在不事先知晓混杂偏差的情况下实现最优估计效率。令人惊讶的是,尽管估计的自适应是可能的,但置信区间长度的自适应在信息论上是不可能的,揭示了推断中存在根本性权衡。
Randomized controlled trials (RCTs) are the gold standard for evaluating the causal effect of a treatment; however, they often have limited sample sizes and sometimes poor generalizability. On the other hand, non-randomized, observational data derived from large administrative databases have massive sample sizes and better generalizability, but they are prone to unmeasured confounding bias. It is thus of considerable interest to reconcile effect estimates obtained from randomized controlled trials and observational studies investigating the same intervention, potentially harvesting the best from both realms. In this paper, we theoretically characterize the potential efficiency gain of integrating observational data into the RCT-based analysis from a minimax point of view. For estimation, we derive the minimax rate of convergence for the mean squared error, and propose a fully adaptive anchored thresholding estimator that attains the optimal rate up to poly-log factors. For inference, we characterize the minimax rate for the length of confidence intervals and show that adaptation (to unknown confounding bias) is in general impossible. A curious phenomenon thus emerges: for estimation, the efficiency gain from data integration can be achieved without prior knowledge on the magnitude of the confounding bias; for inference, the same task becomes information-theoretically impossible in general. We corroborate our theoretical findings using simulations and a real data example from the RCT DUPLICATE initiative [Franklin et al., 2021b].
研究动机与目标
- 解决将高精度但样本量有限的RCT与大规模、可推广但存在混杂的观察性数据相结合以进行因果效应估计的挑战。
- 在极小化极大风险下,理论刻画数据整合带来的效率增益,重点关注估计与推断。
- 研究自适应方法是否能在不事先知晓未测量混杂偏差的情况下实现最优性能。
- 阐明尽管估计中自适应可行,但在置信区间构造中自适应的根本限制。
提出的方法
- 推导使用RCT和观察性数据估计平均处理效应时均方误差的极小化极大收敛速率。
- 提出一种完全自适应的锚定阈值估计器,可在对数因子范围内达到极小化极大速率,且无需事先知晓混杂程度。
- 建立置信区间长度的极小化极大下界,表明对未知混杂偏差的自适应在信息论上是不可能的。
- 采用两样本模型,结合RCT与观察性数据,假设存在未测量混杂的线性模型,并在有界混杂偏差约束下推导极小化极大速率。
- 应用高斯尺度混合与集中不等式,推导在各种偏差情形下置信区间长度的下界。
- 通过模拟和对DUPLICATE RCT倡议的真实世界应用验证理论结果,展示经验性能的提升。
实验结果
研究问题
- RQ1结合RCT与观察性数据时,估计平均处理效应的极小化极大收敛速率是什么?
- RQ2在不事先知晓未测量混杂偏差的情况下,自适应估计器能否实现最优估计性能?
- RQ3当混杂偏差未知时,能否实现构造最优长度置信区间的自适应?
- RQ4观察性数据与RCT数据的样本量比如何影响估计与推断的极小化极大速率?
- RQ5在存在未测量混杂的情况下,估计与推断自适应之间存在何种根本性权衡?
主要发现
- 估计的均方误差极小化极大速率量级为 $ \frac{1}{\sqrt{n_c / \sigma_c^2 + n_o / \sigma_o^2}} $,其中 $ n_c $ 和 $ n_o $ 分别为RCT与观察性数据的样本量。
- 所提出的锚定阈值估计器可在多对数因子范围内达到极小化极大速率,实现无需事先知晓混杂偏差的最优估计。
- 对于推断,置信区间长度的极小化极大速率为 $ \frac{1}{\sqrt{n_c / \sigma_c^2 + n_o / \sigma_o^2}} + \left( \frac{\sigma_c}{\sqrt{n_c}} \land \overline{\Delta} \right) $,其中 $ \overline{\Delta} $ 为未测量混杂的上界。
- 推断中的自适应在信息论上是不可能的:任何估计器都无法在不知晓混杂偏差大小的情况下达到置信区间长度的极小化极大速率。
- 模拟和来自DUPLICATE RCT倡议的真实数据示例表明,与仅使用RCT或仅使用观察性数据的估计器相比,锚定阈值估计器显著降低了均方误差,尤其在混杂较小时效果更明显。
- 当混杂可忽略时,整合观察性数据可带来显著的效率增益,但推断中缺乏自适应性,使得在不知晓偏差的情况下实现最优置信区间覆盖存在根本性限制。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。