[论文解读] Performance of multilabel machine learning models and risk stratification schemas for predicting stroke and bleeding risk in patients with non-valvular atrial fibrillation
本研究评估了多标签机器学习(ML)模型在非瓣膜性心房颤动(NVAF)患者中预测卒中、大出血和死亡的性能,并与标准临床风险评分进行比较。多标签梯度提升机模型在预测大出血(AUC 0.709 vs. 0.522)和死亡(AUC 0.765 vs. 0.606)方面表现优于CHA2DS2-VASc和HAS-BLED,显著提升了区分度,同时识别出血红蛋白和肾功能等新型风险特征。
Appropriate antithrombotic therapy for patients with atrial fibrillation (AF) requires assessment of ischemic stroke and bleeding risks. However, risk stratification schemas such as CHA2DS2-VASc and HAS-BLED have modest predictive capacity for patients with AF. Machine learning (ML) techniques may improve predictive performance and support decision-making for appropriate antithrombotic therapy. We compared the performance of multilabel ML models with the currently used risk scores for predicting outcomes in AF patients. Materials and Methods This was a retrospective cohort study of 9670 patients, mean age 76.9 years, 46% women, who were hospitalized with non-valvular AF, and had 1-year follow-up. The primary outcome was ischemic stroke and major bleeding admission. The secondary outcomes were all-cause death and event-free survival. The discriminant power of ML models was compared with clinical risk scores by the area under the curve (AUC). Risk stratification was assessed using the net reclassification index. Results Multilabel gradient boosting machine provided the best discriminant power for stroke, major bleeding, and death (AUC = 0.685, 0.709, and 0.765 respectively) compared to other ML models. It provided modest performance improvement for stroke compared to CHA2DS2-VASc (AUC = 0.652), but significantly improved major bleeding prediction compared to HAS-BLED (AUC = 0.522). It also had a much greater discriminant power for death compared with CHA2DS2-VASc (AUC = 0.606). Also, models identified additional risk features (such as hemoglobin level, renal function, etc.) for each outcome. Conclusions Multilabel ML models can outperform clinical risk stratification scores for predicting the risk of major bleeding and death in non-valvular AF patients.
研究动机与目标
- 评估多标签机器学习模型在非瓣膜性心房颤动(NVAF)患者中预测多种不良结局的性能。
- 比较ML模型与既定临床风险评分(CHA2DS2-VASc和HAS-BLED)在预测卒中、大出血和死亡方面的区分能力。
- 利用ML识别超出传统风险因素的额外临床特征,以提升结局预测能力。
- 通过净再分类改善和区分度指标评估ML模型的临床实用性。
提出的方法
- 对9,670名具有1年随访的NVAF患者进行回顾性队列研究。
- 使用多标签机器学习模型(包括梯度提升、随机森林和神经网络)同时预测卒中、大出血和死亡。
- 通过受试者工作特征曲线下面积(AUC)评估各结局的模型性能。
- 应用净再分类指数(NRI)评估模型在风险分类方面相对于传统评分的改进程度。
- 通过模型可解释性技术(如特征重要性评分)识别关键预测特征。
- 将模型输出与CHA2DS2-VASc(卒中风险)和HAS-BLED(出血风险)评分作为基准临床工具进行比较。
实验结果
研究问题
- RQ1多标签机器学习模型在预测NVAF患者缺血性卒中方面与CHA2DS2-VASc和HAS-BLED相比表现如何?
- RQ2多标签ML模型能否在HAS-BLED评分基础上进一步改善大出血的预测能力?
- RQ3多标签模型在预测全因死亡率方面是否优于CHA2DS2-VASc?
- RQ4ML模型识别出哪些新型临床特征可作为NVAF患者卒中、出血或死亡的预测因子?
- RQ5与标准临床风险评分相比,ML模型在风险再分类方面改善程度如何?
主要发现
- 多标签梯度提升机模型的区分度最高,其AUC分别为:卒中0.685,大出血0.709,全因死亡0.765。
- 该模型在预测大出血方面显著优于HAS-BLED(AUC 0.709 vs. 0.522),表明区分度有显著提升。
- 在预测死亡方面,其表现也明显优于CHA2DS2-VASc(AUC 0.765 vs. 0.606),表明其预后价值更高。
- 该模型识别出血红蛋白水平和肾功能等临床相关风险特征,作为不良结局的重要预测因子。
- 净再分类改善分析证实,与传统评分相比,ML模型能将患者更准确地重新分类至风险等级。
- 多标签建模实现了多种结局的同步预测,提供了比单一结局模型更全面的风险评估。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。