[论文解读] An explainable Transformer-based deep learning model for the prediction of incident heart failure
本研究提出了一种新颖的、可解释的基于Transformer的深度学习模型,利用来自100,071名英国患者的纵向电子健康记录(EHR)数据,预测新发心力衰竭。该模型在内部验证中AUC-ROC达到0.93,在外部验证中同样达到0.93,优于现有模型,并通过基于扰动的可解释性方法识别出已知和新型风险因素。
Predicting the incidence of complex chronic conditions such as heart failure is challenging. Deep learning models applied to rich electronic health records may improve prediction but remain unexplainable hampering their wider use in medical practice. We developed a novel Transformer deep-learning model for more accurate and yet explainable prediction of incident heart failure involving 100,071 patients from longitudinal linked electronic health records across the UK. On internal 5-fold cross validation and held-out external validation, our model achieved 0.93 and 0.93 area under the receiver operator curve and 0.69 and 0.70 area under the precision-recall curve, respectively and outperformed existing deep learning models. Predictor groups included all community and hospital diagnoses and medications contextualised within the age and calendar year for each patient's clinical encounter. The importance of contextualised medical information was revealed in a number of sensitivity analyses, and our perturbation method provided a way of identifying factors contributing to risk. Many of the identified risk factors were consistent with existing knowledge from clinical and epidemiological research but several new associations were revealed which had not been considered in expert-driven risk prediction models.
研究动机与目标
- 利用丰富的电子健康记录(EHR)数据,通过深度学习提高新发心力衰竭的预测准确性。
- 开发一种既高度准确又内在可解释的模型,以促进临床应用。
- 通过利用EHR中上下文化的医疗信息,识别具有临床意义的风险因素——包括已知和新型风险因素。
- 通过引入基于扰动的解释方法,克服深度学习模型在医疗领域中“黑箱”问题的局限性。
- 通过严格的交叉验证和保留测试,在内部和外部数据集上验证模型的性能。
提出的方法
- 基于Transformer的深度学习架构在纵向EHR数据上进行训练,包括每次临床就诊的诊断、药物、年龄和年份等信息。
- 所有患者数据均通过上下文特征嵌入,为每个医疗事件编码时间与人口统计学背景。
- 应用基于扰动的解释方法,通过系统性地改变输入标记并测量预测变化,评估特征重要性。
- 模型利用多头自注意力机制捕捉时间戳临床事件之间的长程依赖关系。
- 通过5折交叉验证和保留的外部数据集评估性能,以确保稳健性。
- 与现有深度学习基线模型进行比较,以证明其优越的预测性能。
实验结果
研究问题
- RQ1基于Transformer的深度学习模型能否在使用EHR数据时,比现有模型实现更高的新发心力衰竭预测性能?
- RQ2整合上下文化的临床信息(如年龄、年份、诊断序列)在多大程度上能提升预测准确性?
- RQ3该模型能否通过基于扰动的方法提供可靠且人类可解释的预测解释?
- RQ4该模型是否识别出传统专家驱动风险评分中未涵盖的心力衰竭新型风险因素?
- RQ5当在保留的外部数据集上评估时,该模型的性能在不同患者群体中的泛化能力如何?
主要发现
- 在内部5折交叉验证中,模型的受试者工作特征曲线下面积(AUC-ROC)为0.93,在外部保留数据集验证中同样达到0.93。
- 在内部验证中,精确率-召回率曲线下面积(AUC-PR)为0.69,在外部验证中为0.70,表明在数据不平衡情况下表现优异。
- 基于扰动的解释方法成功识别出具有临床相关性的风险因素,包括已确立的和此前未被认识的关联。
- 敏感性分析证实,上下文化的医疗信息(如年龄和年份)对提升预测准确性具有重要意义。
- 发现了若干在专家驱动风险预测模型中此前未被考虑的新型风险关联。
- 在AUC-ROC和AUC-PR指标上,该模型均优于现有深度学习基线模型,展现出更优越的预测能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。