[论文解读] Predicting risk of late age-related macular degeneration using deep learning
本研究开发了一种深度学习模型,将卷积神经网络与生存分析相结合,利用彩色眼底照片预测个体进展为晚期年龄相关性黄斑变性(AMD)的风险。该模型在来自AREDS和AREDS2的3,298名参与者数据上进行训练,五年C统计量达到86.4(95%置信区间:86.2–86.6),显著优于现有临床标准,并在独立队列中展现出高度泛化能力。
By 2040, age-related macular degeneration (AMD) will affect approximately 288 million people worldwide. Identifying individuals at high risk of progression to late AMD, the sight-threatening stage, is critical for clinical actions, including medical interventions and timely monitoring. Although deep learning has shown promise in diagnosing/screening AMD using color fundus photographs, it remains difficult to predict individuals' risks of late AMD accurately. For both tasks, these initial deep learning attempts have remained largely unvalidated in independent cohorts. Here, we demonstrate how deep learning and survival analysis can predict the probability of progression to late AMD using 3,298 participants (over 80,000 images) from the Age-Related Eye Disease Studies AREDS and AREDS2, the largest longitudinal clinical trials in AMD. When validated against an independent test dataset of 601 participants, our model achieved high prognostic accuracy (five-year C-statistic 86.4 (95% confidence interval 86.2-86.6)) that substantially exceeded that of retinal specialists using two existing clinical standards (81.3 (81.1-81.5) and 82.0 (81.8-82.3), respectively). Interestingly, our approach offers additional strengths over the existing clinical standards in AMD prognosis (e.g., risk ascertainment above 50%) and is likely to be highly generalizable, given the breadth of training data from 82 US retinal specialty clinics. Indeed, during external validation through training on AREDS and testing on AREDS2 as an independent cohort, our model retained substantially higher prognostic accuracy than existing clinical standards. These results highlight the potential of deep learning systems to enhance clinical decision-making in AMD patients.
研究动机与目标
- 提高对疾病进展至晚期AMD(致盲阶段)的预测准确性。
- 开发一种将图像分类与生存分析相结合的深度学习框架,用于个体水平的风险预测。
- 在独立数据集(包括AREDS2)上验证模型,评估其在训练数据之外的泛化能力。
- 克服现有临床标准的局限性,例如依赖专家分级特征及预测准确性较低的问题。
- 通过更精确识别高风险个体,实现更早的临床干预。
提出的方法
- 采用两阶段深度学习框架:首先,卷积神经网络(CNN)从彩色眼底照片中提取特征;其次,将这些特征输入Cox比例风险模型进行生存分析。
- 模型在AREDS和AREDS2试验的3,298名参与者数据上进行训练,输入包括双侧眼底图像、年龄、吸烟状态和基因型数据。
- 使用R语言中的'glmnet'包进行特征选择,以识别最具有预测力的协变量并纳入生存模型。
- 通过200次自举重采样估计95%置信区间,使用五年C统计量评估模型性能。
- 利用'keras-vis'包生成显著性图,可视化对预测贡献最大的视网膜区域(例如,玻璃疣、色素异常)。
- 通过在AREDS上训练并在AREDS2上测试,以及反向操作,开展外部验证,以评估模型在不同数据集间的鲁棒性。
实验结果
研究问题
- RQ1在纵向眼底图像上训练的深度学习模型,是否能比现有临床标准更准确地预测个体进展为晚期AMD?
- RQ2将深度特征与生存分析相结合,是否能相比传统风险评分系统提高预后准确性?
- RQ3当在独立队列(如AREDS2)上验证时,该深度学习模型的泛化能力如何?
- RQ4根据模型,哪些视网膜特征(例如,玻璃疣、色素异常)对晚期AMD进展最具预测性?
- RQ5当模型在某一数据集上训练并在另一数据集上测试时,是否能保持高性能,从而证明其在不同临床机构间的鲁棒性?
主要发现
- 该深度学习模型五年C统计量达到86.4(95%置信区间:86.2–86.6),显著优于简化严重度量表(SSS)的81.3(95%置信区间:81.1–81.5)。
- 该模型还优于在线风险计算器(C统计量为82.0,95%置信区间:81.8–82.3),表明其具有更高的预测准确性。
- 在外部验证中,当在AREDS上训练并在AREDS2上测试时,模型的预后准确性仍显著高于现有标准。
- 由于模型在来自82家美国眼科专科诊所的多样化临床环境数据上进行训练,因此展现出高度泛化能力。
- 显著性图显示,模型重点关注临床相关特征,如玻璃疣和色素异常,从而增强了可解释性。
- Brier评分证实了预测概率的校准性得到改善,表明预测风险与实际结果更加一致。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。