[论文解读] Two-phase analysis and study design for survival models with error-prone exposures
该论文将均值评分法扩展至具有测量误差暴露的两阶段生存分析,利用未验证协变量的辅助数据校正测量误差。提出一种自适应抽样设计,在成本约束下通过最小化方差来提高估计效率,在模拟和真实数据中均表现出接近“理想”性能,尤其在最优试点抽样和分层分配下效果更佳。
Increasingly, medical research is dependent on data collected for non-research purposes, such as electronic health records data (EHR). EHR data and other large databases can be prone to measurement error in key exposures, and unadjusted analyses of error-prone data can bias study results. Validating a subset of records is a cost-effective way of gaining information on the error structure, which in turn can be used to adjust analyses for this error and improve inference. We extend the mean score method for the two-phase analysis of discrete-time survival models, which uses the unvalidated covariates as auxiliary variables that act as surrogates for the unobserved true exposures. This method relies on a two-phase sampling design and an estimation approach that preserves the consistency of complete case regression parameter estimates in the validated subset, with increased precision leveraged from the auxiliary data. Furthermore, we develop optimal sampling strategies which minimize the variance of the mean score estimator for a target exposure under a fixed cost constraint. We consider the setting where an internal pilot is necessary for the optimal design so that the phase two sample is split into a pilot and an adaptive optimal sample. Through simulations and data example, we evaluate efficiency gains of the mean score estimator using the derived optimal validation design compared to balanced and simple random sampling for the phase two sample. We also empirically explore efficiency gains that the proposed discrete optimal design can provide for the Cox proportional hazards model in the setting of a continuous-time survival outcome.
研究动机与目标
- 解决在所有受试者均无法获得金标准暴露数据时,生存结局中的测量误差问题。
- 将此前仅用于离散时间生存分析和二值结局的均值评分法,扩展至具有测量误差暴露的连续时间生存模型。
- 在固定成本约束下,开发一种最小化目标回归参数方差的最优两阶段抽样设计。
- 提出一种基于试点阶段估计无需参数的自适应设计,以实现最优分层分配。
- 在模拟和真实世界数据中,评估所提设计相较于简单随机抽样和平衡抽样的效率增益。
提出的方法
- 采用两阶段抽样设计:第一阶段在所有受试者中收集易于获得的测量误差暴露数据;第二阶段对子样本进行验证,以估计误差结构。
- 应用均值评分法,利用测量误差暴露作为辅助变量,在离散时间比例风险模型中校正回归参数估计。
- 通过Neyman分配推导最优抽样比例,在固定验证样本量下最小化目标参数的方差。
- 引入多波抽样策略:试点阶段估计无需参数(如分层特定均值和方差),从而实现在主阶段的自适应分配。
- 采用改进的自适应设计,根据试点数据调整抽样策略,相较于平衡或简单随机抽样性能更优。
- 将结果扩展至使用Cox比例风险模型的连续时间生存分析,证明在相同设计框架下可实现效率增益。
实验结果
研究问题
- RQ1均值评分法能否有效扩展至具有测量误差暴露的离散时间生存模型的两阶段分析?
- RQ2如何在固定成本约束下推导最优抽样比例,以最小化目标回归参数的方差?
- RQ3利用试点阶段估计无需参数对自适应抽样设计效率有何影响?
- RQ4所提出的自适应设计相较于简单随机抽样和平衡抽样,在方差减少和估计效率方面表现如何?
- RQ5所推导的最优设计在使用Cox模型进行连续时间生存分析时,能在多大程度上提高效率?
主要发现
- 在国家威尔姆斯肿瘤研究数据中,所提出的自适应抽样设计结合均值评分估计器,相较于平衡抽样和简单随机抽样,方差显著降低达14%。
- 包含此前分析中被排除的间歇性删失个体(n = 158)后,精度进一步提高,尤其对MS-BAL和MS-A估计器更为明显。
- 该方法在阶段二样本量增加时,相对于完整案例分析及其他抽样策略,表现出一致的效率增益。
- 在所有模拟设置和数据实例中,自适应设计均优于简单随机抽样和平衡抽样,展现出稳健性和实际应用价值。
- 该方法成功扩展至使用Cox模型的连续时间生存结局分析,在I型删失和随机右删失下均观察到效率增益。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。