Skip to main content
QUICK REVIEW

[论文解读] Disability prediction in multiple sclerosis using performance outcome measures and demographic data

Subhrajit Roy, Diana Mincu|arXiv (Cornell University)|Apr 8, 2022
Multiple Sclerosis Research Studies被引用 4
一句话总结

本研究首次(据我们所知)证明,仅使用表现结果测量(POMs)和人口统计学数据,机器学习模型即可准确预测多发性硬化症(MS)的残疾进展,从而在临床和基于智能手机的环境中实现可靠、可扩展且低成本的监测。该方法在多种数据集和模型类型中均表现出强劲的预测性能,且在不同人口统计子组中结果一致,对特征删减也具有鲁棒性。

ABSTRACT

Literature on machine learning for multiple sclerosis has primarily focused on the use of neuroimaging data such as magnetic resonance imaging and clinical laboratory tests for disease identification. However, studies have shown that these modalities are not consistent with disease activity such as symptoms or disease progression. Furthermore, the cost of collecting data from these modalities is high, leading to scarce evaluations. In this work, we used multi-dimensional, affordable, physical and smartphone-based performance outcome measures (POM) in conjunction with demographic data to predict multiple sclerosis disease progression. We performed a rigorous benchmarking exercise on two datasets and present results across 13 clinically actionable prediction endpoints and 6 machine learning models. To the best of our knowledge, our results are the first to show that it is possible to predict disease progression using POMs and demographic data in the context of both clinical trials and smartphone-base studies by using two datasets. Moreover, we investigate our models to understand the impact of different POMs and demographics on model performance through feature ablation studies. We also show that model performance is similar across different demographic subgroups (based on age and sex). To enable this work, we developed an end-to-end reusable pre-processing and machine learning framework which allows quicker experimentation over disparate MS datasets.

研究动机与目标

  • 探究仅使用表现结果测量(POMs)和人口统计学数据,是否可在不依赖神经影像学或实验室检查的情况下,预测多发性硬化症(MS)的残疾进展。
  • 在两个公开的MS数据集(MSOAC和Floodlight)上对多种机器学习模型进行基准测试。
  • 通过年龄和性别等人口统计子组评估模型性能,以确保公平性和泛化能力。
  • 开发一个可复用的端到端预处理与建模框架,以实现对异构MS数据集的高效基准测试。
  • 通过删减研究理解各POM和人口统计学特征对预测性能的相对贡献。

提出的方法

  • 本研究使用多维、带时间戳的POMs(包括步行、平衡、认知和精细运动功能评估)数据,通过临床随访(MSOAC)或智能手机应用(Floodlight)收集,同时结合人口统计学数据(年龄、性别)。
  • 采用标准化、可复用的预处理流程,将多样化的MS数据集转换为统一格式,以实现标签的一致创建和指标的统一计算。
  • 评估六种机器学习模型:XGBoost、LightGBM、CatBoost、DenseNet、TCN和Transformer,涵盖13个临床可操作的预测终点(如EDSS >3、EDSS >5、EDSS严重程度类别)和两个时间范围(6–12个月和12–24个月)。
  • 使用AU-PRC评估模型性能,并按年龄和性别进行子组分析,以评估公平性和鲁棒性。
  • 通过系统性删减POMs和人口统计学特征,量化其对预测性能的个体贡献。
  • 通过TCN和Transformer模型探索时序建模,以评估POM中序列模式对长期预测的价值。

实验结果

研究问题

  • RQ1仅使用POMs和人口统计学数据,是否可在临床试验和基于智能手机的环境中,以具有临床意义的准确度预测MS残疾进展?
  • RQ2不同机器学习模型在多种MS数据集和预测时间范围内的表现如何?
  • RQ3模型性能是否在不同人口统计子组(如年龄和性别)间存在显著差异,表明可能存在偏差或公平性问题?
  • RQ4哪些POMs和人口统计学特征对预测性能贡献最大?人口统计学数据是否对实现高精度至关重要?
  • RQ5可复用的端到端框架是否能够实现对异构MS数据集的高效基准测试,并推动未来MS进展预测模型的开发?

主要发现

  • 本研究在MSOAC和Floodlight两个数据集中均实现了强劲的预测性能,其中TCN和Transformer模型因有效建模时序模式,在长期预测中表现更优。
  • 模型性能在人口统计子组中保持一致:男性和女性的AU-PRC值相似,50–70岁年龄组的表现与全数据集相当或略低,尽管队列以复发缓解型为主。
  • 最年轻年龄组(30岁以下)的AU-PRC下降最为显著,提示预测早发性MS进展可能存在挑战。
  • 不包含人口统计学数据的POMs表现与完整特征集相当,引发对人口统计学数据在预测性能中必要性的质疑,并凸显隐私权衡问题。
  • 特征删减研究显示,某些POMs(特别是与运动功能和认知相关的)对模型性能的贡献显著高于其他POMs。
  • 所提出的端到端框架可实现可靠的数据集摄入、可扩展的标签创建和一致的指标计算,支持在异构MS数据集间快速实验。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。