[Paper Review] Disease Trajectory Maps
The paper introduces the Disease Trajectory Map (DTM), a probabilistic model that learns low-dimensional, interpretable representations of sparse and irregularly sampled clinical time series from electronic health records. It uses stochastic variational inference to scale to large datasets and demonstrates on scleroderma data that DTM-learned embeddings reveal biologically meaningful subpopulations and significantly associate with key clinical outcomes, outperforming LMM and FPCA baselines in detecting clinically relevant trajectory-outcome links.
Medical researchers are coming to appreciate that many diseases are in fact complex, heterogeneous syndromes composed of subpopulations that express different variants of a related complication. Time series data extracted from individual electronic health records (EHR) offer an exciting new way to study subtle differences in the way these diseases progress over time. In this paper, we focus on answering two questions that can be asked using these databases of time series. First, we want to understand whether there are individuals with similar disease trajectories and whether there are a small number of degrees of freedom that account for differences in trajectories across the population. Second, we want to understand how important clinical outcomes are associated with disease trajectories. To answer these questions, we propose the Disease Trajectory Map (DTM), a novel probabilistic model that learns low-dimensional representations of sparse and irregularly sampled time series. We propose a stochastic variational inference algorithm for learning the DTM that allows the model to scale to large modern medical datasets. To demonstrate the DTM, we analyze data collected on patients with the complex autoimmune disease, scleroderma. We find that DTM learns meaningful representations of disease trajectories and that the representations are significantly associated with important clinical outcomes.
Motivation & Objective
- To identify latent subpopulations of patients with similar disease progression trajectories using clinical time series from electronic health records.
- To determine whether a small number of latent dimensions can capture the major sources of variation in disease trajectories across heterogeneous populations.
- To investigate associations between learned trajectory representations and clinically significant outcomes such as organ damage or mortality.
- To develop a scalable, probabilistic model that handles irregularly sampled and sparse biomedical time series data.
- To outperform existing methods (LMM, FPCA) in uncovering biologically meaningful and clinically relevant trajectory patterns.
Proposed method
- The DTM is a probabilistic model that learns low-dimensional (2D or 3D) latent representations of individual clinical marker trajectories.
- It formulates the problem as a latent variable model where each patient’s trajectory is generated from a low-dimensional embedding and a flexible, non-linear function of time.
- The model uses a warped Gaussian process prior over random effects, enabling it to capture complex, non-linear trajectory shapes.
- A stochastic variational inference algorithm is developed to scale the inference to large EHR datasets with thousands of patients.
- The model is trained using held-out log-likelihood estimation via 10-fold cross-validation to ensure generalization performance.
- Associations between representations and clinical outcomes are tested using kernel density estimator-based hypothesis tests.
Experimental results
Research questions
- RQ1Are there distinct, biologically meaningful subpopulations of patients with similar disease progression trajectories in complex autoimmune diseases?
- RQ2Can a low-dimensional latent space capture the primary sources of variation in irregular and sparse clinical time series data?
- RQ3Do patients with specific adverse clinical outcomes (e.g., pulmonary hypertension, organ damage) exhibit significantly different trajectory representations?
- RQ4How do the DTM’s learned representations compare to those from LMM and FPCA in detecting clinically relevant patterns?
- RQ5Can the DTM uncover previously unknown associations between disease trajectory shapes and outcomes?
Key findings
- The DTM learned 2D and 3D representations that reveal distinct subpopulations of scleroderma patients consistent with prior clinical findings.
- The DTM achieved comparable generalization performance to LMM and FPCA, with a held-out log-likelihood of -13.25 ± 1.38 on the PFVC marker and -3.32 ± 0.06 on the TSS marker.
- The DTM detected a statistically significant association between trajectory representations and pulmonary arterial hypertension (PAH), with a p-value of 0.002, which was not detected by FPCA or LMM.
- For interstitial lung disease, all three models rejected the null hypothesis (p < 0.001), confirming expected associations.
- The DTM uniquely revealed a significant link between TSS trajectory representations and ulcers/gangrene (p = 0.009), which was missed by LMM and FPCA.
- The DTM’s representations showed stronger statistical evidence for associations with myositis and pulmonary hypertension than baseline models, particularly in the presence of fibrosis-related risk factors.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.