[论文解读] Principal stratification analysis using principal scores
本文提出一种基于主分层得分(principal scores)的主分层框架,即给定协变量下潜在分层的条件概率,以实现在具有中间变量的随机化实验中稳健、无需模型假设的因果推断。通过基于这些得分对样本加权,该方法在不假设正态性或参数化结果模型的前提下实现有效的因果推断,同时为关键的不可检验假设(如忽略性与单调性)提供敏感性分析。
Practitioners are interested in not only the average causal effect of the treatment on the outcome but also the underlying causal mechanism in the presence of an intermediate variable between the treatment and outcome. However, in many cases we cannot randomize the intermediate variable, resulting in sample selection problems even in randomized experiments. Therefore, we view randomized experiments with intermediate variables as semi-observational studies. In parallel with the analysis of observational studies, we provide a theoretical foundation for conducting objective causal inference with an intermediate variable under the principal stratification framework, with principal strata defined as the joint potential values of the intermediate variable. Our strategy constructs weighted samples based on principal scores, defined as the conditional probabilities of the latent principal strata given covariates, without access to any outcome data. This principal stratification analysis yields robust causal inference without relying on any model assumptions on the outcome distributions. We also propose approaches to conducting sensitivity analysis for violations of the ignorability and monotonicity assumptions, the very crucial but untestable identification assumptions in our theory. When the assumptions required by the classical instrumental variable analysis cannot be justified by background knowledge or cannot be made because of scientific questions of interest, our strategy serves as a useful alternative tool to deal with intermediate variables. We illustrate our methodologies by using two real data examples, and find scientifically meaningful conclusions.
研究动机与目标
- 解决在中间变量影响结果但无法随机化的随机化实验中因果推断的挑战。
- 克服经典工具变量方法依赖于不可检验的排除限制和单调性假设的局限性。
- 开发一种无需模型、客观的方法,仅基于治疗前协变量估计主分层内的因果效应。
- 为忽略性与单调性等关键但不可检验的假设提供敏感性分析工具,这些假设在主分层中至关重要。
- 在传统方法因缺乏排除限制或强参数假设而失效的场景中,实现实际的因果推断。
提出的方法
- 将主分层定义为在处理与对照下中间变量的联合潜在值,并将其视为潜在的治疗前协变量。
- 引入主分层得分为给定观测协变量下属于各主分层的条件概率,且在无需结果数据的情况下进行估计。
- 利用主分层得分作为逆概率权重,构建主分层内平均因果效应的加权估计量。
- 在主分层忽略性假设下,推导加权估计量的大样本性质与标准误。
- 通过引入一个敏感性参数 ξ 来对单调性进行敏感性分析,以量化对单调性的偏离程度。
- 利用基于主分层得分的加权方法,在不依赖结果模型假设的前提下,保持协变量平衡并提高估计效率。
实验结果
研究问题
- RQ1当无法假设排除限制与单调性时,如何在具有中间变量的随机化实验中实现有效的因果推断?
- RQ2我们能否在不假设主分层内结果分布的参数模型的前提下,实现稳健的因果推断?
- RQ3治疗前协变量在确保忽略性与提高主分层中估计效率方面发挥什么作用?
- RQ4如何评估因果估计对单调性与主分层忽略性假设违反的敏感性?
- RQ5主分层得分能否用于构建稳定、无需模型的估计量,以避免在有限混合模型中出现似然函数无界的問題?
主要发现
- 所提出的基于主分层得分的加权方法可在不需对结果分布做参数假设的前提下,实现主分层内平均因果效应的有效、无需模型的因果推断。
- 对单调性的敏感性分析表明,估计的幸存者平均因果效应(ACE_ss)在敏感性参数 ξ 的合理取值范围内保持稳健,其区间估计始终覆盖零。
- 在 SWOG 前列腺癌试验中,治疗组的“幸存者”分层主分层得分估计为 0.496,对照组为 0.389,表明对生存的影响存在中等但非单调的治疗效应。
- 该方法通过平衡检查确认了治疗组间协变量的平衡,支持忽略性假设的合理性。
- 该方法避免了正态混合模型带来的数值不稳定性与模型敏感性问题,为在参数假设存疑时提供了更可靠的替代方案。
- 理论结果表明,包含预测性协变量可提高估计效率,强化了在实验设计中收集丰富治疗前协变量的重要性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。