Skip to main content
QUICK REVIEW

[论文解读] Handling missing values in healthcare data: A systematic review of deep learning-based imputation techniques

M. Liu, Siqi Li|arXiv (Cornell University)|Oct 15, 2022
Machine Learning in Healthcare参考文献 92被引用 4
一句话总结

本篇系统性综述评估了用于医疗数据的基于深度学习(DL)的填补技术,分析了模型架构、填补策略及数据类型。研究发现,DL方法在准确性方面优于非DL方法,尤其在时间序列和多模态数据中表现更优,其中‘整合’策略——即填补与下游任务联合优化——在复杂数据场景中展现出显著有效性。

ABSTRACT

Objective: The proper handling of missing values is critical to delivering reliable estimates and decisions, especially in high-stakes fields such as clinical research. The increasing diversity and complexity of data have led many researchers to develop deep learning (DL)-based imputation techniques. We conducted a systematic review to evaluate the use of these techniques, with a particular focus on data types, aiming to assist healthcare researchers from various disciplines in dealing with missing values. Methods: We searched five databases (MEDLINE, Web of Science, Embase, CINAHL, and Scopus) for articles published prior to August 2021 that applied DL-based models to imputation. We assessed selected publications from four perspectives: health data types, model backbone (i.e., main architecture), imputation strategies, and comparison with non-DL-based methods. Based on data types, we created an evidence map to illustrate the adoption of DL models. Results: We included 64 articles, of which tabular static (26.6%, 17/64) and temporal data (37.5%, 24/64) were the most frequently investigated. We found that model backbone(s) differed among data types as well as the imputation strategy. The "integrated" strategy, that is, the imputation task being solved concurrently with downstream tasks, was popular for tabular temporal (50%, 12/24) and multi-modal data (71.4%, 5/7), but limited for other data types. Moreover, DL-based imputation methods yielded better imputation accuracy in most studies, compared with non-DL-based methods. Conclusion: DL-based imputation models can be customized based on data type, addressing the corresponding missing patterns, and its associated "integrated" strategy can enhance the efficacy of imputation, especially in scenarios where data is complex. Future research may focus on the portability and fairness of DL-based models for healthcare data imputation.

研究动机与目标

  • 评估当前基于深度学习的医疗数据填补技术的研究现状。
  • 识别DL填补研究中使用最广泛的的数据类型和模型架构。
  • 评估不同填补策略的有效性,特别是将填补与下游任务结合的‘整合’方法。
  • 从准确性和可靠性角度,对比DL填补方法与传统非DL技术(如均值填补、k-NN、多重填补)的性能差异。
  • 提供一个结构化的证据地图,以指导研究人员根据数据类型和缺失模式选择合适的DL模型。

提出的方法

  • 在五个数据库(MEDLINE、Web of Science、Embase、CINAHL和Scopus)中开展系统性文献检索,截止至2021年8月。
  • 筛选出64项应用深度学习模型对医疗数据中的缺失值进行填补的研究。
  • 根据数据类型(如表格静态数据、时间序列数据、多模态数据)、模型主干架构(如自编码器、生成对抗网络、Transformer)和填补策略(如独立填补、整合填补)对研究进行分类。
  • 通过将DL填补结果与非DL基线方法(如均值填补、k-最近邻、多重填补)进行比较,评估模型性能。
  • 生成证据地图,可视化不同医疗数据类型中DL填补技术的分布情况。
  • 分析‘整合’策略的普及程度与有效性,即填补过程与下游预测任务联合优化。

实验结果

研究问题

  • RQ1在医疗数据中,哪些数据类型最常被基于深度学习的填补方法所针对?
  • RQ2不同深度学习模型架构(如自编码器、生成对抗网络、Transformer)在各类医疗数据类型中的表现如何?
  • RQ3在复杂数据环境中,‘整合’填补策略相较于独立填补策略的相对有效性如何?
  • RQ4DL填补技术在准确性方面相较于非DL方法(如均值填补或k-NN)表现如何?
  • RQ5在不同临床数据模态和缺失模式下,DL填补技术的采用模式呈现出何种趋势?

主要发现

  • 时间序列数据(37.5%,24/64)和表格静态数据(26.6%,17/64)是在DL填补研究中最常被研究的数据类型。
  • ‘整合’填补策略——即填补与下游任务联合训练——在时间序列数据研究中占50%,在多模态数据研究中占71.4%。
  • 在所审查的大多数研究中,基于DL的填补方法在填补准确性方面始终优于非DL方法。
  • 模型主干架构在不同数据类型间存在显著差异,表明模型架构的选择应根据特定数据结构和缺失模式进行定制。
  • 证据地图揭示了DL模型在复杂、高维医疗数据中应用的持续增长趋势,尤其在纵向和多模态场景中更为明显。
  • 尽管性能表现优异,但当前研究对模型可移植性及在多样化患者群体中的公平性相关挑战仍关注不足。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。