Skip to main content
QUICK REVIEW

[Paper Review] Electronic health record phenotyping improves detection and screening of type 2 diabetes in the general United States population: A cross-sectional, unselected, retrospective study

Ariana Anderson, Wesley T. Kerr|arXiv (Cornell University)|Jan 10, 2015
Diabetes Management and Education35 references22 citations
TL;DR

This study demonstrates that electronic health record (EHR) phenotyping significantly improves type 2 diabetes detection in the U.S. general population by leveraging comprehensive EHR data—including diagnoses, medications, and clinical variables—via multivariate logistic regression and random forests. The full EHR model outperformed conventional risk models using only BMI, age, gender, and lifestyle factors (p<0.001), identifying novel associations such as migraines and cardiac dysrhythmias as negatively linked to diabetes.

ABSTRACT

Objectives: In the United States, 25% of people with type 2 diabetes are undiagnosed. Conventional screening models use limited demographic information to assess risk. We evaluated whether electronic health record (EHR) phenotyping could improve diabetes screening, even when records are incomplete and data are not recorded systematically across patients and practice locations. Methods: In this cross-sectional, retrospective study, data from 9,948 US patients between 2009 and 2012 were used to develop a pre-screening tool to predict current type 2 diabetes, using multivariate logistic regression. We compared (1) a full EHR model containing prescribed medications, diagnoses, and traditional predictive information, (2) a restricted EHR model where medication information was removed, and (3) a conventional model containing only traditional predictive information (BMI, age, gender, hypertensive and smoking status). We additionally used a random-forests classification model to judge whether including additional EHR information could increase the ability to detect patients with Type 2 diabetes on new patient samples. Results: Using a patient's full or restricted EHR to detect diabetes was superior to using basic covariates alone (p&lt;0.001). The random forests model replicated on out-of-bag data. Migraines and cardiac dysrhythmias were negatively associated with type 2 diabetes, while acute bronchitis and herpes zoster were positively associated, among other factors. Conclusions: EHR phenotyping resulted in markedly superior detection of type 2 diabetes in a general US population, could increase the efficiency and accuracy of disease screening, and are capable of picking up signals in real-world records.

Motivation & Objective

  • To improve detection of undiagnosed type 2 diabetes in the U.S. population, where 25% remain undiagnosed.
  • To evaluate whether EHR-derived phenotypes can enhance screening accuracy despite incomplete or inconsistent data recording.
  • To compare the performance of EHR-based models against conventional risk models relying only on demographic and basic clinical factors.
  • To identify novel clinical associations linked to type 2 diabetes using real-world EHR data.
  • To validate the robustness of EHR phenotyping using out-of-sample prediction with random forests.

Proposed method

  • Conducted a cross-sectional, retrospective analysis of EHR data from 9,948 U.S. patients (2009–2012).
  • Developed a multivariate logistic regression model using full EHR data (diagnoses, medications, traditional predictors) to predict current type 2 diabetes status.
  • Constructed a restricted EHR model excluding medication data to assess its impact on predictive performance.
  • Built a conventional model using only basic covariates: BMI, age, gender, hypertension, and smoking status.
  • Applied a random forests classification model to evaluate predictive performance on out-of-bag data and detect new signal associations.
  • Used receiver operating characteristic (ROC) analysis to compare model performance across all three models.

Experimental results

Research questions

  • RQ1Can EHR phenotyping significantly improve the detection of type 2 diabetes compared to conventional screening models?
  • RQ2How does the inclusion of medication data affect the predictive accuracy of EHR-based diabetes screening models?
  • RQ3What novel clinical associations (beyond traditional risk factors) are predictive of type 2 diabetes in real-world EHR data?
  • RQ4Can EHR phenotyping maintain high performance even when data are incomplete or inconsistently recorded across patients and settings?
  • RQ5To what extent can random forests models replicate and generalize findings from logistic regression in EHR-based phenotyping?

Key findings

  • The full EHR model significantly outperformed the conventional model in detecting type 2 diabetes (p < 0.001), demonstrating superior discrimination.
  • Even the restricted EHR model (excluding medications) showed significantly better performance than the conventional model (p < 0.001), indicating diagnostic and clinical data alone improve detection.
  • Random forests models successfully replicated results on out-of-bag data, confirming model robustness and generalizability.
  • Migraines and cardiac dysrhythmias were negatively associated with type 2 diabetes, suggesting potential protective or inverse clinical signals.
  • Acute bronchitis and herpes zoster were positively associated with type 2 diabetes, indicating possible comorbid or predictive associations.
  • EHR phenotyping effectively extracted meaningful biological signals from real-world, unstructured, and incomplete clinical records, enhancing screening efficiency and accuracy.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.