Skip to main content
QUICK REVIEW

[Paper Review] Prediction of Coronary Heart Disease Using Routine Blood Tests

Menɡ Ninɡ, Peng Zhang|arXiv (Cornell University)|Sep 12, 2018
Cardiovascular Function and Risk FactorsMedicine6 references4 citations
TL;DR

This study develops a two-layer Gradient Boosting Decision Tree (GBDT) model using routine blood test data to predict coronary heart disease (CHD) risk, achieving 86% sensitivity in identifying CHD patients. The model leverages 15,000 blood test records to classify health, CHD, and other diseases, revealing clear patterns in blood markers associated with CHD for clinical use.

ABSTRACT

Background --The objective of this study was to examine the association of routine blood test results with coronary heart disease (CHD) risk, to incorporate them into coronary prediction models and to compare the discrimination properties of this approach with other prediction functions. Methods and Results --This work was designed as a retrospective, single-center study of a hospital-based cohort. The 5060 CHD patients (2365 men and 2695 women) were 1 to 97 years old at baseline with 8 years (2009-2017) of medical records, 5051 health check-ups and 5075 cases of other diseases. We developed a two-layer Gradient Boosting Decision Tree(GBDT) model based on routine blood data to predict the risk of coronary heart disease, which could identify 86% of people with coronary heart disease. We built a dataset with 15,000 routine blood tests results. Using this dataset, we trained the two-layer GBDT model to classify healthy status, coronary heart disease and other diseases. As a result of the classification after machine learning, we found that the sensitivity of detecting the health data was approximately 93% for all data, and the sensitivity of detecting CHD was 93% for disease data that included coronary heart disease. On this basis, we further visualized the correlation between routine blood results and related data items, and there was an obvious pattern in health and coronary heart disease in all data presentations, which can be used for clinical reference. Finally, we briefly analyzed the results above from the perspective of pathophysiology. Conclusions --Routine blood data provides more information about CHD than what we already know through the correlation between test results and related data items. A simple coronary disease prediction model was developed using a GBDT algorithm, which will allow physicians to predict CHD risk in patients without overt CHD.

Motivation & Objective

  • To investigate the association between routine blood test results and coronary heart disease (CHD) risk.
  • To develop a machine learning model that improves CHD risk prediction using standard laboratory data.
  • To compare the performance of the proposed model with existing prediction functions.
  • To visualize correlations between blood test results and CHD for clinical interpretability.
  • To provide a pathophysiological interpretation of the identified biomarker patterns.

Proposed method

  • A retrospective, single-center study was conducted using 8 years of medical records (2009–2017) from a hospital cohort.
  • A dataset of 15,000 routine blood test results was compiled, including 5,060 CHD cases, 5,051 healthy check-ups, and 5,075 other disease cases.
  • A two-layer Gradient Boosting Decision Tree (GBDT) model was trained to classify three categories: healthy, CHD, and other diseases.
  • The model used feature importance and correlation analysis to identify key blood markers linked to CHD.
  • Model performance was evaluated using sensitivity, specificity, and classification accuracy across the three classes.
  • Pathophysiological interpretations were provided based on the identified blood test patterns.

Experimental results

Research questions

  • RQ1Can routine blood test results improve the prediction of coronary heart disease risk beyond conventional methods?
  • RQ2What is the performance of a GBDT-based model in classifying CHD using only standard blood test data?
  • RQ3Which blood test parameters show the strongest association with CHD in the dataset?
  • RQ4Can visualizations of blood test patterns enhance clinical understanding of CHD risk?
  • RQ5How do the identified biomarker patterns align with known pathophysiological mechanisms of CHD?

Key findings

  • The GBDT model achieved 86% sensitivity in identifying patients with coronary heart disease using routine blood test data.
  • The model demonstrated 93% sensitivity in detecting healthy individuals across the full dataset.
  • Sensitivity for CHD detection was 93% when analyzing disease-related data, indicating strong classification performance.
  • Clear, distinct patterns emerged in blood test results between healthy individuals and those with CHD, supporting clinical interpretability.
  • The model successfully differentiated CHD from other diseases, suggesting specificity in identifying CHD-related biomarker profiles.
  • Pathophysiological analysis revealed plausible biological links between the identified blood markers and coronary heart disease mechanisms.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.