[Paper Review] Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques
This systematic review evaluates deep learning (DL)-based imputation techniques for healthcare data, analyzing model architectures, imputation strategies, and data types. It finds that DL methods outperform non-DL approaches in accuracy, especially for temporal and multi-modal data, with 'integrated' strategies—where imputation is co-optimized with downstream tasks—showing strong efficacy in complex data scenarios.
Objective: The proper handling of missing values is critical to delivering reliable estimates and decisions, especially in high-stakes fields such as clinical research. The increasing diversity and complexity of data have led many researchers to develop deep learning (DL)-based imputation techniques. We conducted a systematic review to evaluate the use of these techniques, with a particular focus on data types, aiming to assist healthcare researchers from various disciplines in dealing with missing values. Methods: We searched five databases (MEDLINE, Web of Science, Embase, CINAHL, and Scopus) for articles published prior to August 2021 that applied DL-based models to imputation. We assessed selected publications from four perspectives: health data types, model backbone (i.e., main architecture), imputation strategies, and comparison with non-DL-based methods. Based on data types, we created an evidence map to illustrate the adoption of DL models. Results: We included 64 articles, of which tabular static (26.6%, 17/64) and temporal data (37.5%, 24/64) were the most frequently investigated. We found that model backbone(s) differed among data types as well as the imputation strategy. The "integrated" strategy, that is, the imputation task being solved concurrently with downstream tasks, was popular for tabular temporal (50%, 12/24) and multi-modal data (71.4%, 5/7), but limited for other data types. Moreover, DL-based imputation methods yielded better imputation accuracy in most studies, compared with non-DL-based methods. Conclusion: DL-based imputation models can be customized based on data type, addressing the corresponding missing patterns, and its associated "integrated" strategy can enhance the efficacy of imputation, especially in scenarios where data is complex. Future research may focus on the portability and fairness of DL-based models for healthcare data imputation.
Motivation & Objective
- To evaluate the current state of deep learning-based imputation techniques in healthcare data.
- To identify the most commonly used data types and model architectures in DL imputation research.
- To assess the effectiveness of different imputation strategies, particularly the 'integrated' approach that combines imputation with downstream tasks.
- To compare DL-based imputation methods with traditional non-DL techniques in terms of accuracy and reliability.
- To provide a structured evidence map to guide researchers in selecting appropriate DL models based on data type and missingness patterns.
Proposed method
- Conducted a systematic literature search across five databases: MEDLINE, Web of Science, Embase, CINAHL, and Scopus, up to August 2021.
- Selected 64 studies that applied deep learning models to impute missing values in healthcare data.
- Classified studies based on data type (e.g., tabular static, temporal, multi-modal), model backbone (e.g., autoencoders, GANs, transformers), and imputation strategy (e.g., standalone, integrated).
- Evaluated model performance by comparing DL-based imputation results with non-DL baselines such as mean imputation, k-NN, and multiple imputation.
- Generated an evidence map to visualize the distribution of DL imputation techniques across different healthcare data types.
- Analyzed the prevalence and effectiveness of the 'integrated' strategy, where imputation is jointly optimized with downstream prediction tasks.
Experimental results
Research questions
- RQ1Which data types in healthcare are most frequently targeted by deep learning-based imputation methods?
- RQ2How do different deep learning model architectures (e.g., autoencoders, GANs, transformers) perform across various healthcare data types?
- RQ3What is the relative effectiveness of 'integrated' imputation strategies compared to standalone imputation in complex data settings?
- RQ4How do DL-based imputation techniques compare in accuracy to non-DL-based methods like mean imputation or k-NN?
- RQ5What patterns emerge in the adoption of DL imputation techniques across different clinical data modalities and missingness patterns?
Key findings
- Temporal data (37.5%, 24/64) and tabular static data (26.6%, 17/64) were the most commonly studied data types in DL imputation research.
- The 'integrated' imputation strategy—where imputation is co-trained with downstream tasks—was used in 50% of studies on temporal data and 71.4% of studies on multi-modal data.
- DL-based imputation methods consistently outperformed non-DL-based methods in terms of imputation accuracy across most studies reviewed.
- Model backbones varied significantly by data type, indicating that architecture choice should be tailored to the specific data structure and missingness pattern.
- The evidence map revealed a growing trend in applying DL models to complex, high-dimensional healthcare data, particularly in longitudinal and multi-modal settings.
- Despite strong performance, challenges related to model portability and fairness across diverse patient populations remain under-investigated in current research.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.