[Paper Review] Disability prediction in multiple sclerosis using performance outcome measures and demographic data
This study demonstrates, for the first time to our knowledge, that machine learning models can accurately predict multiple sclerosis (MS) disability progression using only performance outcome measures (POMs) and demographic data—enabling reliable, scalable, and low-cost monitoring in both clinical and smartphone-based settings. The approach achieves strong predictive performance across diverse datasets and model types, with consistent results across demographic subgroups and robustness to feature ablation.
Literature on machine learning for multiple sclerosis has primarily focused on the use of neuroimaging data such as magnetic resonance imaging and clinical laboratory tests for disease identification. However, studies have shown that these modalities are not consistent with disease activity such as symptoms or disease progression. Furthermore, the cost of collecting data from these modalities is high, leading to scarce evaluations. In this work, we used multi-dimensional, affordable, physical and smartphone-based performance outcome measures (POM) in conjunction with demographic data to predict multiple sclerosis disease progression. We performed a rigorous benchmarking exercise on two datasets and present results across 13 clinically actionable prediction endpoints and 6 machine learning models. To the best of our knowledge, our results are the first to show that it is possible to predict disease progression using POMs and demographic data in the context of both clinical trials and smartphone-base studies by using two datasets. Moreover, we investigate our models to understand the impact of different POMs and demographics on model performance through feature ablation studies. We also show that model performance is similar across different demographic subgroups (based on age and sex). To enable this work, we developed an end-to-end reusable pre-processing and machine learning framework which allows quicker experimentation over disparate MS datasets.
Motivation & Objective
- To investigate whether performance outcome measures (POMs) and demographic data can predict MS disability progression without relying on neuroimaging or lab tests.
- To benchmark multiple machine learning models across two publicly available MS datasets: MSOAC and Floodlight.
- To evaluate model performance across demographic subgroups (age and sex) to ensure fairness and generalizability.
- To develop a reusable, end-to-end pre-processing and modeling framework for efficient benchmarking across disparate MS datasets.
- To understand the relative contribution of individual POMs and demographic features through ablation studies.
Proposed method
- The study uses multi-dimensional, time-stamped POMs—including walking, balance, cognition, and dexterity assessments—collected via clinical visits (MSOAC) or smartphone apps (Floodlight), along with demographic data (age, sex).
- A standardized, reusable pre-processing pipeline ingests diverse MS datasets into a common format, enabling consistent label creation and metric computation.
- Six machine learning models are evaluated: XGBoost, LightGBM, CatBoost, DenseNet, TCN, and Transformer, across 13 clinically actionable prediction endpoints (e.g., EDSS >3, EDSS >5, EDSS severity category) and two time horizons (6–12 months and 12–24 months).
- Model performance is assessed using AU-PRC, with subgroup analysis by age and sex to evaluate fairness and robustness.
- Feature ablation studies systematically remove POMs and demographic features to quantify their individual contributions to predictive performance.
- Temporal modeling is explored via TCN and Transformer models to assess the value of sequential patterns in POMs for long-term prediction.
Experimental results
Research questions
- RQ1Can POMs and demographic data alone predict MS disability progression with clinically meaningful accuracy in both clinical trial and smartphone-based settings?
- RQ2How do different machine learning models perform across diverse MS datasets and prediction horizons?
- RQ3Does model performance vary significantly across demographic subgroups (e.g., age and sex), indicating potential bias or fairness issues?
- RQ4Which POMs and demographic features contribute most to predictive performance, and is demographic data essential for high accuracy?
- RQ5Can a reusable, end-to-end framework enable efficient benchmarking and future development of MS progression prediction models across heterogeneous datasets?
Key findings
- The study achieves strong predictive performance across both MSOAC and Floodlight datasets, with TCN and Transformer models showing superior performance on long-term predictions due to effective modeling of temporal patterns.
- Model performance remains consistent across demographic subgroups: males and females show similar AU-PRC values, and performance for age groups 50–70 is on par or slightly lower than the full dataset, despite the cohort being predominantly relapsing-remitting.
- The youngest age group (under 30) shows the most significant drop in AU-PRC, suggesting potential challenges in predicting early-onset MS progression.
- POMs without demographic data perform on par with the full feature set, raising questions about the necessity of demographic data for predictive performance and highlighting privacy trade-offs.
- Feature ablation reveals that certain POMs—particularly those related to motor function and cognition—contribute more significantly to model performance than others.
- The proposed end-to-end framework enables reliable dataset ingestion, scalable label creation, and consistent metric computation, supporting rapid experimentation across disparate MS datasets.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.