Skip to main content
QUICK REVIEW

[Paper Review] Personalized Imputation in metric spaces via conformal prediction: Applications in Predicting Diabetes Development with Continuous Glucose Monitoring Information

Marcos Matabuena, Carla Díaz‐Louzao|arXiv (Cornell University)|Mar 26, 2024
Gene expression and cancer classificationBiochemistry, Genetics and Molecular Biology3 citations
TL;DR

This paper proposes a novel two-step framework for personalized imputation of missing continuous glucose monitoring (CGM) data in metric spaces using conformal prediction, enabling improved prediction of diabetes onset. By modeling glucose profiles as probability distributions (glucodensities) in 2-Wasserstein space and applying personalized conformal prediction, the method boosts predictive accuracy by over 10% compared to traditional models.

ABSTRACT

The challenge of handling missing data is widespread in modern data analysis, particularly during the preprocessing phase and in various inferential modeling tasks. Although numerous algorithms exist for imputing missing data, the assessment of imputation quality at the patient level often lacks personalized statistical approaches. Moreover, there is a scarcity of imputation methods for metric space based statistical objects. The aim of this paper is to introduce a novel two-step framework that comprises: (i) a imputation methods for statistical objects taking values in metrics spaces, and (ii) a criterion for personalizing imputation using conformal inference techniques. This work is motivated by the need to impute distributional functional representations of continuous glucose monitoring (CGM) data within the context of a longitudinal study on diabetes, where a significant fraction of patients do not have available CGM profiles. The importance of these methods is illustrated by evaluating the effectiveness of CGM data as new digital biomarkers to predict the time to diabetes onset in healthy populations. To address these scientific challenges, we propose: (i) a new regression algorithm for missing responses; (ii) novel conformal prediction algorithms tailored for metric spaces with a focus on density responses within the 2-Wasserstein geometry; (iii) a broadly applicable personalized imputation method criterion, designed to enhance both of the aforementioned strategies, yet valid across any statistical model and data structure. Our findings reveal that incorporating CGM data into diabetes time-to-event analysis, augmented with a novel personalization phase of imputation, significantly enhances predictive accuracy by over ten percent compared to traditional predictive models for time to diabetes.

Motivation & Objective

  • To address the lack of personalized, statistically rigorous imputation methods for missing functional and distributional data in biomedical studies.
  • To develop a method for imputing missing continuous glucose monitoring (CGM) profiles represented as probability distributions in metric spaces.
  • To integrate high-resolution glucodensity data as digital biomarkers for predicting time-to-diabetes onset in non-diabetic populations.
  • To provide a generalizable, model-agnostic framework for personalized imputation that enhances predictive performance in longitudinal health studies.
  • To validate the method in a real-world longitudinal cohort study (AEGIS) where only a subset of participants had complete CGM data.

Proposed method

  • Proposes a weighted least squares estimator for linear models in metric spaces, enabling regression on statistical objects like probability distributions.
  • Introduces a novel conformal prediction algorithm tailored for metric spaces, specifically using the 2-Wasserstein distance to quantify uncertainty in distributional responses.
  • Applies conditional Fréchet means as imputation targets within bounded metric spaces, ensuring consistency and robustness under distributional data.
  • Employs a personalized imputation criterion based on conformal prediction that adapts to individual-level uncertainty, enhancing reliability and calibration.
  • Uses glucodensity representations derived from raw CGM data to capture full temporal glucose dynamics, replacing summary statistics.
  • Validates the method through a two-step study design: imputation of missing CGM profiles followed by time-to-event analysis using C-index and AUC.
(a) Glucodensity profiles from raw CGM data for a diabetic and non diabetic individual.
(a) Glucodensity profiles from raw CGM data for a diabetic and non diabetic individual.

Experimental results

Research questions

  • RQ1Can personalized imputation of missing CGM data in metric spaces improve the prediction of time-to-diabetes onset in non-diabetic populations?
  • RQ2How does incorporating distributional representations of glucose profiles (glucodensities) enhance predictive performance compared to traditional biomarkers?
  • RQ3To what extent does conformal prediction in 2-Wasserstein space improve the reliability and calibration of imputed distributional responses?
  • RQ4Can a general-purpose, model-agnostic imputation criterion be developed that enhances predictive accuracy across diverse data structures and statistical models?
  • RQ5What is the impact of personalized uncertainty quantification on the C-index and AUC in time-to-event prediction models using high-resolution glucose data?

Key findings

  • The proposed method improves predictive accuracy for time-to-diabetes onset by over 10% compared to traditional models relying only on scalar biomarkers.
  • The personalized conformal prediction framework achieves a C-index of 0.90 at a radius of 110, encompassing 62 subjects with imputed CGM data, demonstrating strong calibration and coverage.
  • The Area Under the Curve (AUC) over time consistently exceeds that of traditional CGM risk assessments, indicating superior dynamic predictive performance.
  • The conditional Fréchet mean with prediction bands effectively captures individualized glucose dynamics, with uncertainty increasing with larger radii in the conformal prediction framework.
  • The method significantly enhances model performance even when only a subset of the cohort (580 out of 1,516) has complete CGM data, validating its utility in cost-constrained longitudinal studies.
  • The C-score for the model incorporating personalized imputed CGM data reaches 0.805, representing a substantial improvement over models without functional data.
(b) Glucodensities profiles of all subjects with CGM, separated according to the status of diabetes. Red: individuals with diabetes at baseline. Black: individuals without diabetes at baseline who developed diabetes throughout the study. Grey: individuals free of diabetes at the end of the study.
(b) Glucodensities profiles of all subjects with CGM, separated according to the status of diabetes. Red: individuals with diabetes at baseline. Black: individuals without diabetes at baseline who developed diabetes throughout the study. Grey: individuals free of diabetes at the end of the study.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.