Skip to main content
QUICK REVIEW

[论文解读] A Comparison of Different Methods to Adjust Survival Curves for Confounders

Robin Denz, Renate Klaaßen‐Mielke|arXiv (Cornell University)|Mar 18, 2022
Health Systems, Economic Evaluations, Quality of Life被引用 5
一句话总结

本研究比较了10种在观察性研究中校正混杂因素的生存曲线调整方法,包括逆概率治疗加权法(IPTW)、G-公式、倾向得分匹配、经验似然法以及增广估计量。基于真实数据和蒙特卡洛模拟,研究发现双重稳健的增广估计量(如AIPTW)在各种情景下均表现出最小偏差和良好拟合度,尤其在结果模型或处理模型正确设定时,优于小样本中的标准Cox回归方法。

ABSTRACT

Treatment specific survival curves are an important tool to illustrate the treatment effect in studies with time-to-event outcomes. In non-randomized studies, unadjusted estimates can lead to biased depictions due to confounding. Multiple methods to adjust survival curves for confounders exist. However, it is currently unclear which method is the most appropriate in which situation. Our goal is to compare forms of Inverse Probability of Treatment Weighting, the G-Formula, Propensity Score Matching, Empirical Likelihood Estimation and augmented estimators as well as their pseudo-values based counterparts in different scenarios with a focus on their bias and goodness-of-fit. We provide a short review of all methods and illustrate their usage by contrasting the survival of smokers and non-smokers, using data from the German Epidemiological Trial on Ankle-Brachial-Index. Subsequently, we compare the methods using a Monte-Carlo simulation. We consider scenarios in which correctly or incorrectly specified models for describing the treatment assignment and the time-to-event outcome are used with varying sample sizes. The bias and goodness-of-fit is determined by taking the entire survival curve into account. When used properly, all methods showed no systematic bias in medium to large samples. Cox regression based methods, however, showed systematic bias in small samples. The goodness-of-fit varied greatly between different methods and scenarios. Methods utilizing an outcome model were more efficient than other techniques, while augmented estimators using an additional treatment assignment model were unbiased when either model was correct with a goodness-of-fit comparable to other methods. These doubly-robust methods have important advantages in every considered scenario.

研究动机与目标

  • 评估并比较在非随机化研究中对生存曲线进行调整的方法性能,以校正未调整的Kaplan-Meier估计所受的混杂偏倚。
  • 评估在不同样本量和模型设定(正确/错误指定的处理模型与结果模型)下,各方法的偏差和拟合优度。
  • 识别在现实观察性数据条件下估计反事实生存曲线最稳健且高效的估计方法。
  • 通过在真实数据和模拟情景下的性能评估,为方法选择提供实用指导。

提出的方法

  • 本研究评估了10种方法:IPTW、G-公式、倾向得分匹配、经验似然法(EL)、增广估计量(AIPTW)及其伪值对应方法。
  • 使用德国踝臂指数流行病学试验的真实世界数据,说明在比较吸烟者与非吸烟者生存差异时的方法应用。
  • 开展一项全面的蒙特卡洛模拟研究,共重复2000次,样本量分别为100、500和1000,对处理分配和事件时间结果模型均采用正确和错误指定的设定。
  • 在整个生存曲线上计算偏差和均方误差(MSE),而非固定时间点,以评估整体曲线的准确性。
  • 通过评估覆盖概率和平均置信区间宽度,评估区间估计性能。
  • 使用时间上的综合偏差和MSE对各方法进行比较,特别关注在模型设定错误时的稳健性。

实验结果

研究问题

  • RQ1在不同样本量和模型设定下,哪种方法能产生对反事实生存曲线的最小偏差估计?
  • RQ2当结果模型或处理模型之一被错误指定时,增广估计量(如AIPTW)与Cox回归或IPTW等标准方法相比表现如何?
  • RQ3样本量对不同调整技术的偏差和拟合优度有何影响?
  • RQ4基于伪值的方法与原始方法相比,在偏差和效率方面表现如何?
  • RQ5在现实观察性研究设定中,哪种方法在偏差减少、精度和覆盖概率之间提供了最佳平衡?

主要发现

  • 当模型正确设定时,所有方法在中等及以上样本量下均无系统性偏差,但Cox回归方法在小样本中表现出显著偏差。
  • 利用结果模型的方法(如G-公式、EL、AIPTW)比仅依赖处理分配模型的方法效率更高。
  • 增广估计量(如AIPTW)在结果模型或处理模型其中之一正确设定时,可实现无偏估计,表现出‘双重稳健’特性。
  • 不同方法和情景下的拟合优度差异显著,AIPTW和AIPTW PV在偏差和MSE方面表现更优,尤其在小样本中。
  • Kaplan-Meier估计在控制组和治疗组中均表现出显著偏差(最高达±0.019),尤其在小样本中,证实了校正的必要性。
  • 绝大多数方法的95%置信区间覆盖概率接近名义水平,尽管部分方法(如IPTW KM)在小样本中置信区间略宽。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。