[论文解读] Machine Learning-Assisted Recurrence Prediction for Early-Stage Non-Small-Cell Lung Cancer Patients
本研究开发并评估了基于表格和图结构的机器学习模型,利用来自1,387名西班牙患者的临床数据,预测早期非小细胞肺癌(NSCLC)患者的复发情况。表现最佳的模型为基于表格数据的随机森林模型,在10折交叉验证中达到76%的准确率,SHAP与基于实例的解释方法增强了模型的可解释性,有助于临床决策支持。
Background: Stratifying cancer patients according to risk of relapse can personalize their care. In this work, we provide an answer to the following research question: How to utilize machine learning to estimate probability of relapse in early-stage non-small-cell lung cancer patients? Methods: For predicting relapse in 1,387 early-stage (I-II), non-small-cell lung cancer (NSCLC) patients from the Spanish Lung Cancer Group data (65.7 average age, 24.8% females, 75.2% males) we train tabular and graph machine learning models. We generate automatic explanations for the predictions of such models. For models trained on tabular data, we adopt SHAP local explanations to gauge how each patient feature contributes to the predicted outcome. We explain graph machine learning predictions with an example-based method that highlights influential past patients. Results: Machine learning models trained on tabular data exhibit a 76% accuracy for the Random Forest model at predicting relapse evaluated with a 10-fold cross-validation (model was trained 10 times with different independent sets of patients in test, train and validation sets, the reported metrics are averaged over these 10 test sets). Graph machine learning reaches 68% accuracy over a 200-patient, held-out test set, calibrated on a held-out set of 100 patients. Conclusions: Our results show that machine learning models trained on tabular and graph data can enable objective, personalised and reproducible prediction of relapse and therefore, disease outcome in patients with early-stage NSCLC. With further prospective and multisite validation, and additional radiological and molecular data, this prognostic model could potentially serve as a predictive decision support tool for deciding the use of adjuvant treatments in early-stage lung cancer. Keywords: Non-Small-Cell Lung Cancer, Tumor Recurrence Prediction, Machine Learning
研究动机与目标
- 在传统分期之外,提升对早期NSCLC患者个性化复发风险预测的能力。
- 解决尽管进行了根治性切除术,仍有30–55%的患者术后复发率较高的问题。
- 利用临床数据开发客观、可复现且可解释的机器学习模型,用于复发预测。
- 整合表格数据与图结构学习方法,以捕捉患者层面和人群层面的复杂关系。
- 通过提供个体化的复发风险评分,支持临床决策,为辅助治疗提供依据。
提出的方法
- 在西班牙肺癌集团胸腺肿瘤登记库提供的1,387名早期NSCLC患者的临床特征上,训练了基于表格数据的机器学习模型(包括随机森林)。
- 应用SHAP(SHapley Additive exPlanations)生成表格模型预测的局部、特征级别的解释。
- 开发了一种图结构机器学习模型,将患者和临床概念表示为知识图谱,以捕捉长距离依赖关系。
- 采用基于实例的解释方法,检索并对比具有影响力的既往患者,以解释图模型的预测结果。
- 对表格模型进行10折交叉验证,对图模型使用200名患者的保留测试集,并在100名患者的集合上进行校准。
- 聚焦于二元复发预测(是/否),而非生存时间预测,以减少非癌症相关死亡的混杂影响。
实验结果
研究问题
- RQ1与标准分期相比,基于表格临床数据训练的机器学习模型是否能提高对早期NSCLC患者复发预测的准确性?
- RQ2整合异构临床数据与人群层面数据的图结构机器学习模型,是否能超越表格模型,进一步提升复发预测能力?
- RQ3如何通过可解释人工智能技术使模型预测更具可解释性并可临床应用?
- RQ4在真实世界、回顾性的早期NSCLC患者队列中,可解释模型在复发预测方面的表现如何?
- RQ5这些模型能否支持低风险与高风险患者在辅助治疗方面的个性化决策?
主要发现
- 基于表格数据的随机森林模型在10折交叉验证中达到76%的准确率,表现出较强的预测性能。
- 图结构机器学习模型在200名患者的保留测试集中达到68%的准确率,尽管性能低于表格模型,但仍显示出良好前景。
- SHAP解释揭示了影响个体复发风险预测的关键临床特征,提升了模型的透明度。
- 基于实例的解释方法成功检索并突出显示了具有影响力的既往患者,增强了图模型输出的可解释性。
- 模型基于单一国家队列(西班牙)训练,表明其在泛化性方面可能存在人口学限制。
- 未来整合影像学与分子数据有望进一步提升模型性能与临床实用性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。