Skip to main content
QUICK REVIEW

[论文解读] Bayesian nonparametric inference for the covariate-adjusted ROC curve

Vanda Inácio, María Xosé Rodríguez‐Álvarez|Edinburgh Research Explorer|May 30, 2018
Bayesian Methods and Mixture Models参考文献 8被引用 11
一句话总结

本文提出了一种基于B样条和依附狄利克雷过程混合模型结合贝叶斯重抽样法的贝叶斯非参数方法,用于估计校正协变量后的受试者工作特征(ROC)曲线。该方法在调整年龄、性别等协变量的同时,实现了对诊断试验准确性的稳健、灵活推断,在模拟研究中展示了在复杂情景下对真实ROC曲线的准确恢复能力以及有效的不确定性量化。

ABSTRACT

Accurate diagnosis of disease is of fundamental importance in clinical practice and medical research. Before a medical diagnostic test is routinely used in practice, its ability to distinguish between diseased and nondiseased states must be rigorously assessed through statistical analysis. The receiver operating characteristic (ROC) curve is the most popular used tool for evaluating the discriminatory ability of continuous-outcome diagnostic tests. It has been acknowledged that several factors (e.g., subject-specific characteristics, such as age and/or gender) can affect the test's accuracy beyond disease status. Recently, the covariate-adjusted ROC curve has been proposed and successfully applied as a global summary measure of diagnostic accuracy that takes covariate information into account. We motivate the use of the covariate-adjusted ROC curve and develop a highly robust model based on a combination of B-splines dependent Dirichlet process mixture models and the Bayesian bootstrap. Multiple simulation studies demonstrate the ability of our model to successfully recover the true covariate-adjusted ROC curve and to produce valid inferences in a variety of complex scenarios. Our methods are motivated by and applied to an endocrine study where the main goal is to assess the accuracy of the body mass index, adjusted for age and gender, for predicting clusters of cardiovascular disease risk factors. The R-package AROC, implementing our proposed methods, is provided.

研究动机与目标

  • 为解决标准ROC分析忽略受试者特异性协变量(如年龄和性别)的问题,这些协变量可能显著影响诊断试验准确性。
  • 开发一种灵活且稳健的统计模型,允许在不假设参数分布形式的前提下进行协变量特定的ROC曲线估计。
  • 通过将协变量信息直接整合到ROC曲线估计中,避免分层或过度简化的假设,从而提高推断准确性。
  • 提供一个完全贝叶斯的非参数框架,通过后验分布和可信区间实现不确定性量化。
  • 在真实模拟情景和一项关于BMI预测心血管风险因素的内分泌学实际研究中,展示该方法的性能。

提出的方法

  • 采用具有B样条的依附狄利克雷过程混合模型,灵活建模给定协变量的诊断试验结果的条件分布。
  • 使用贝叶斯重抽样法非参数估计患病组和非患病组中试验结果的条件累积分布函数。
  • 通过整合条件分布的后验预测分布,将协变量校正的ROC曲线建模为协变量的函数。
  • 在狄利克雷过程的混合测度上施加非参数先验,以实现灵活且由数据驱动的结果分布聚类。
  • 通过B样条基函数实现节点位置和光滑性控制,将内部节点数(K)视为超参数。
  • 使用WAIC和LPML准则进行模型比较与选择,涵盖不同节点数和先验设定的情形。
Figure 2: True (solid black line) and average value of 100 simulated datasets (dashed lines) of the posterior mean (for the Bayesian estimators) of the covariate adjusted ROC curve/pooled ROC curve for each of the scenarios and sample size $(n_{\bar{D}},n_{D})=(200,200)$ . The shaded area are bands
Figure 2: True (solid black line) and average value of 100 simulated datasets (dashed lines) of the posterior mean (for the Bayesian estimators) of the covariate adjusted ROC curve/pooled ROC curve for each of the scenarios and sample size $(n_{\bar{D}},n_{D})=(200,200)$ . The shaded area are bands

实验结果

研究问题

  • RQ1在复杂、非正态且异方差的数据结构下,贝叶斯非参数模型能否准确恢复真实的协变量校正ROC曲线?
  • RQ2在预测心血管风险因素时,纳入年龄和性别等协变量如何影响BMI诊断准确性的估计?
  • RQ3在多种模拟情景下,与参数或半参数替代方法相比,该方法在覆盖性和偏差方面的表现如何?
  • RQ4不同的先验设定(如浓度参数、尺度超参数)如何影响后验推断和模型拟合?
  • RQ5该方法能否通过可信区间在一系列协变量取值范围内有效覆盖真实ROC曲线,实现有效的不确定性量化?

主要发现

  • 所提出的贝叶斯非参数方法在所有六个模拟情景中均成功恢复了真实的协变量校正ROC曲线,即使在复杂、非正态且异方差的数据结构下亦然。
  • 从该方法构建的后验可信区间在所有情景和假阳性率水平(0.1和0.3)下均实现了约95%的覆盖率,表明不确定性量化有效。
  • 在ROC曲线估计的覆盖率和均方误差方面,具有四个内部节点(K=4)的模型始终优于无内部节点(K=0)的模型。
  • 在BMI与心血管风险的真实世界应用中,该方法揭示了诊断准确性在年龄和性别之间存在显著差异,老年男性和年轻女性的AUC更高。
  • 基于WAIC和LPML的模型比较表明,具有更多内部节点(K=4)的模型拟合效果更优,随着K增加,WAIC上升而LPML下降,表明灵活性提升带来更好的模型拟合。
  • 实现该方法的R包 AROC 已公开发布,支持可重现且易于访问的真实诊断试验评估应用。
(a) Age-specific AUC
(a) Age-specific AUC

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。