[论文解读] Identifying Cancer Patients at Risk for Heart Failure Using Machine Learning Methods
本研究利用电子健康记录(EHR)数据开发了一种机器学习模型,用于在心脏毒性治疗前预测癌症患者的心力衰竭风险。基于143,199名患者的梯度提升模型,AUC达到0.9077,表现出高敏感性和特异性,显示出在癌症治疗相关心毒性早期干预方面的潜力。
Cardiotoxicity related to cancer therapies has become a serious issue, diminishing cancer treatment outcomes and quality of life. Early detection of cancer patients at risk for cardiotoxicity before cardiotoxic treatments and providing preventive measures are potential solutions to improve cancer patients's quality of life. This study focuses on predicting the development of heart failure in cancer patients after cancer diagnoses using historical electronic health record (EHR) data. We examined four machine learning algorithms using 143,199 cancer patients from the University of Florida Health (UF Health) Integrated Data Repository (IDR). We identified a total number of 1,958 qualified cases and matched them to 15,488 controls by gender, age, race, and major cancer type. Two feature encoding strategies were compared to encode variables as machine learning features. The gradient boosting (GB) based model achieved the best AUC score of 0.9077 (with a sensitivity of 0.8520 and a specificity of 0.8138), outperforming other machine learning methods. We also looked into the subgroup of cancer patients with exposure to chemotherapy drugs and observed a lower specificity score (0.7089). The experimental results show that machine learning methods are able to capture clinical factors that are known to be associated with heart failure and that it is feasible to use machine learning methods to identify cancer patients at risk for cancer therapy-related heart failure.
研究动机与目标
- 识别在癌症诊断后心力衰竭风险升高的癌症患者。
- 利用历史电子健康记录(EHR)数据,通过机器学习进行风险预测。
- 比较多种机器学习算法在预测癌症患者心力衰竭发生方面的表现。
- 评估特征编码策略对模型性能的改善作用。
- 评估模型在亚组中的表现,特别是接受过化疗的患者。
提出的方法
- 本研究使用来自UF Health综合数据仓库的143,199名癌症患者的回顾性队列数据。
- 将1,958例心力衰竭病例按性别、年龄、种族和主要癌症类型以1:8的比例与对照组匹配。
- 训练并评估了四种机器学习算法:梯度提升、随机森林、逻辑回归和XGBoost。
- 应用了两种特征编码策略:独热编码和目标编码,以表示分类变量。
- 通过AUC、敏感性和特异性衡量模型性能,并采用交叉验证以确保稳健性。
- 对接受过化疗的患者进行了亚组分析,以评估模型的泛化能力。
实验结果
研究问题
- RQ1基于EHR数据训练的机器学习模型能否准确预测癌症患者的心力衰竭发生?
- RQ2在预测癌症患者心力衰竭风险方面,哪种机器学习算法表现最佳?
- RQ3不同的特征编码策略如何影响该临床预测任务中的模型性能?
- RQ4模型在接受化疗的患者中的表现是否显著不同?
- RQ5通过早期识别高风险患者,能否通过预防性干预改善临床结局?
主要发现
- 梯度提升(GB)模型的AUC最高,达到0.9077,优于其他算法。
- GB模型的敏感性为0.8520,特异性为0.8138,表明其具有强大的预测能力。
- 在接受化疗的患者亚组中,特异性下降至0.7089,表明该群体中性能有所降低。
- 模型成功捕捉到了与心力衰竭相关的已知临床危险因素,验证了其临床相关性。
- 特征编码策略显著影响模型性能,其中目标编码在某些情境下表现更优。
- 研究结果支持利用EHR数据进行机器学习,以主动识别存在治疗相关心力衰竭风险的癌症患者。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。