Skip to main content
QUICK REVIEW

[Paper Review] Unsupervised shape and motion analysis of 3822 cardiac 4D MRIs of UK Biobank

Qiao Zheng, Hervé Delingette|arXiv (Cornell University)|Feb 15, 2019
Radiomics and Machine Learning in Medical Imaging33 references4 citations
TL;DR

This study proposes an unsupervised clustering approach on 3,822 cardiac 4D MRIs from the UK Biobank using deep learning-derived shape and motion features. Using Gaussian mixture modeling after feature selection, two distinct clusters emerged that likely represent pathological cardiac phenotypes, validated via classification models and dimensionality reduction, demonstrating the potential of unsupervised analysis on large unlabeled medical imaging datasets.

ABSTRACT

We perform unsupervised analysis of image-derived shape and motion features extracted from 3822 cardiac 4D MRIs of the UK Biobank. First, with a feature extraction method previously published based on deep learning models, we extract from each case 9 feature values characterizing both the cardiac shape and motion. Second, a feature selection is performed to remove highly correlated feature pairs. Third, clustering is carried out using a Gaussian mixture model on the selected features. After analysis, we identify two small clusters which probably correspond to two pathological categories. Further confirmation using a trained classification model and dimensionality reduction tools is carried out to support this discovery. Moreover, we examine the differences between the other large clusters and compare our measures with the ground-truth.

Motivation & Objective

  • To perform unsupervised analysis on a large, unlabeled cardiac MRI dataset from the UK Biobank.
  • To identify potential pathological subtypes without relying on expert-labeled data.
  • To evaluate the effectiveness of clustering in detecting abnormal cardiac phenotypes from image-derived features.
  • To validate the discovered clusters using additional machine learning tools such as classification models and dimensionality reduction.
  • To demonstrate the feasibility of unsupervised learning for cardiac pathology detection in big medical imaging data.

Proposed method

  • Extracted 9 shape and motion features per case using a previously published deep learning-based feature extraction pipeline.
  • Performed feature selection to eliminate highly correlated feature pairs, improving clustering stability.
  • Applied Gaussian mixture modeling (GMM) to cluster cases based on the selected features without supervision.
  • Validated cluster interpretations using a trained classification model to assess pathological likelihood.
  • Utilized dimensionality reduction techniques (PCA and t-SNE) to visualize and confirm cluster separation and structure.
  • Compared automatic pipeline-derived measures with ground-truth values from the InlineVF algorithm to assess reliability and identify outliers.

Experimental results

Research questions

  • RQ1Can unsupervised clustering of image-derived shape and motion features detect pathological subtypes in a large, unlabeled cardiac MRI dataset?
  • RQ2Do the identified clusters correspond to clinically meaningful cardiac pathologies?
  • RQ3How do the feature-based clustering results compare to ground-truth ejection fraction and volume measurements?
  • RQ4To what extent do dimensionality reduction and classification models support the pathological interpretation of the clusters?
  • RQ5Are the discrepancies between automatic pipeline estimates and ground-truth values due to data quality issues?

Key findings

  • Two small clusters were identified that likely correspond to pathological cardiac categories, based on clustering and validation with a trained classifier.
  • The t-SNE and PCA visualizations confirmed distinct separation of these two clusters from the larger, presumably healthy, groups.
  • The automatic pipeline's LVC volume estimates at end-diastole were on average 7.0% lower than ground-truth values, while at end-systole, they were 40.8% lower.
  • Robust linear regression showed strong agreement between automatic pipeline and ground-truth measures, with regression lines closely aligning with the identity line.
  • Outliers in the ground-truth data—likely due to failures in the InlineVF algorithm—were identified as a major source of discrepancy between ground-truth and automatic estimates.
  • The study demonstrates that unsupervised learning on large unlabeled datasets can reveal clinically relevant subtypes without prior labeling.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.