Skip to main content
QUICK REVIEW

[Paper Review] Development and Validation of ML-DQA -- a Machine Learning Data Quality Assurance Framework for Healthcare

Mark Sendak, Gaurav Sirdeshmukh|arXiv (Cornell University)|Aug 4, 2022
Electronic Health Records Systems8 citations
TL;DR

This paper introduces ML-DQA, a machine learning data quality assurance framework designed to standardize and validate real-world healthcare data across diverse clinical projects. Applied across five multi-geographic, multi-condition studies involving 247,536 patients, ML-DQA enabled consistent data quality practices—such as automated rule-based transformations, unified data element mapping, and clinical adjudication—resulting in 2,999 quality checks and 24 reports, with 23.4 data elements per project transformed or removed on average.

ABSTRACT

The approaches by which the machine learning and clinical research communities utilize real world data (RWD), including data captured in the electronic health record (EHR), vary dramatically. While clinical researchers cautiously use RWD for clinical investigations, ML for healthcare teams consume public datasets with minimal scrutiny to develop new algorithms. This study bridges this gap by developing and validating ML-DQA, a data quality assurance framework grounded in RWD best practices. The ML-DQA framework is applied to five ML projects across two geographies, different medical conditions, and different cohorts. A total of 2,999 quality checks and 24 quality reports were generated on RWD gathered on 247,536 patients across the five projects. Five generalizable practices emerge: all projects used a similar method to group redundant data element representations; all projects used automated utilities to build diagnosis and medication data elements; all projects used a common library of rules-based transformations; all projects used a unified approach to assign data quality checks to data elements; and all projects used a similar approach to clinical adjudication. An average of 5.8 individuals, including clinicians, data scientists, and trainees, were involved in implementing ML-DQA for each project and an average of 23.4 data elements per project were either transformed or removed in response to ML-DQA. This study demonstrates the importance role of ML-DQA in healthcare projects and provides teams a framework to conduct these essential activities.

Motivation & Objective

  • Address the inconsistency in data quality practices between clinical researchers and ML for healthcare teams using real-world data (RWD).
  • Bridge the gap between cautious clinical RWD use and minimal scrutiny in public dataset consumption for ML model development.
  • Develop a scalable, reusable framework grounded in RWD best practices to ensure data quality across diverse healthcare ML projects.
  • Standardize data quality processes across multiple projects, geographies, and medical conditions to improve reproducibility and reliability.
  • Demonstrate the framework's feasibility and impact through real-world implementation in five distinct ML projects.

Proposed method

  • Design a modular framework, ML-DQA, integrating standardized data quality checks across data elements in electronic health records (EHRs).
  • Implement a common library of rule-based transformations to normalize diagnosis and medication data elements across projects.
  • Establish a unified method for assigning data quality checks to specific data elements based on clinical and technical criteria.
  • Apply a consistent approach to clinical adjudication of data quality issues involving multidisciplinary teams including clinicians, data scientists, and trainees.
  • Group redundant data element representations using a standardized mapping strategy to reduce data heterogeneity.
  • Automate data quality reporting through integrated utilities, generating 24 quality reports across five projects.

Experimental results

Research questions

  • RQ1How can a standardized data quality assurance framework improve consistency and reliability in healthcare ML projects using real-world data?
  • RQ2To what extent can a unified framework be applied across diverse clinical conditions, geographies, and data sources in real-world healthcare settings?
  • RQ3What are the key reusable practices that emerge when applying a centralized data quality framework across multiple ML projects?
  • RQ4How does involving multidisciplinary teams (clinicians, data scientists, trainees) impact the implementation and effectiveness of data quality assurance?
  • RQ5What measurable improvements in data quality and transformation can be achieved using ML-DQA across multiple projects?

Key findings

  • ML-DQA was successfully applied across five ML projects involving 247,536 patients across two geographies and multiple medical conditions.
  • A total of 2,999 data quality checks were executed and 24 quality reports were generated, demonstrating systematic data validation.
  • An average of 23.4 data elements per project were either transformed or removed in response to data quality issues identified by ML-DQA.
  • Five generalizable data quality practices emerged: standardized data element grouping, automated diagnosis/medication mapping, shared rule-based transformation library, unified check assignment, and consistent clinical adjudication.
  • On average, 5.8 individuals per project—including clinicians, data scientists, and trainees—were involved in implementing ML-DQA, indicating strong interdisciplinary collaboration.
  • The framework enabled consistent, reproducible data quality practices across heterogeneous projects, significantly reducing data inconsistency and improving model reliability.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.