Skip to main content
QUICK REVIEW

[Paper Review] Assessing Intra-class Diversity and Quality of Synthetically Generated Images in a Biomedical and Non-biomedical Setting

Muhammad Saad, Mubashir Husain Rehmani|arXiv (Cornell University)|Jul 23, 2023
Advanced Image Processing TechniquesComputer Science3 citations
TL;DR

This study evaluates intra-class diversity and quality of synthetically generated images using MS-SSIM, Cosine Distance (CD), and Frechet Inception Distance (FID) in biomedical (X-ray, OCT) and non-biomedical (Fashion MNIST) settings. Results show significant metric variance across imaging modalities due to distinct image features, with sample size having no significant impact on scores, highlighting modality-dependent GAN performance and evaluation challenges.

ABSTRACT

In biomedical image analysis, data imbalance is common across several imaging modalities. Data augmentation is one of the key solutions in addressing this limitation. Generative Adversarial Networks (GANs) are increasingly being relied upon for data augmentation tasks. Biomedical image features are sensitive to evaluating the efficacy of synthetic images. These features can have a significant impact on metric scores when evaluating synthetic images across different biomedical imaging modalities. Synthetically generated images can be evaluated by comparing the diversity and quality of real images. Multi-scale Structural Similarity Index Measure and Cosine Distance are used to evaluate intra-class diversity, while Frechet Inception Distance is used to evaluate the quality of synthetic images. Assessing these metrics for biomedical and non-biomedical imaging is important to investigate an informed strategy in evaluating the diversity and quality of synthetic images. In this work, an empirical assessment of these metrics is conducted for the Deep Convolutional GAN in a biomedical and non-biomedical setting. The diversity and quality of synthetic images are evaluated using different sample sizes. This research intends to investigate the variance in diversity and quality across biomedical and non-biomedical imaging modalities. Results demonstrate that the metrics scores for diversity and quality vary significantly across biomedical-to-biomedical and biomedical-to-non-biomedical imaging modalities.

Motivation & Objective

  • Investigate the impact of sample size on intra-class diversity and quality metrics for synthetic images in biomedical and non-biomedical domains.
  • Assess inconsistencies in diversity and quality metric scores between two biomedical imaging modalities: X-ray and Optical Coherence Tomography (OCT).
  • Analyze variability in evaluation metric scores when comparing synthetic images across biomedical and non-biomedical imaging domains.
  • Provide empirical insights into the reliability of standard metrics (MS-SSIM, CD, FID) for evaluating GAN-generated images in diverse imaging contexts.
  • Highlight the influence of inherent image features on metric performance to inform better evaluation strategies for synthetic biomedical data.

Proposed method

  • Employed Deep Convolutional GAN (DCGAN) to generate synthetic images across biomedical (X-ray, OCT) and non-biomedical (Fashion MNIST) datasets.
  • Used Multi-scale Structural Similarity Index Measure (MS-SSIM) and Cosine Distance (CD) to quantify intra-class diversity of synthetic images.
  • Applied Frechet Inception Distance (FID) to evaluate the perceptual quality and distributional similarity of synthetic images to real images.
  • Evaluated metrics across varying sample sizes (25%, 50%, 75%, 100%) to assess sensitivity to training data quantity.
  • Compared metric scores between real and synthetic images within and across imaging modalities to detect performance disparities.
  • Analyzed score distributions across different image classes and modalities to identify modality-specific patterns in diversity and quality.

Experimental results

Research questions

  • RQ1How does varying sample size affect the intra-class diversity and quality scores of synthetic images in biomedical and non-biomedical settings?
  • RQ2To what extent do MS-SSIM and CD scores differ between synthetic and real images in X-ray and OCT modalities?
  • RQ3How do FID scores vary across different biomedical and non-biomedical image classes, and what explains these variations?
  • RQ4What is the impact of inherent image features (e.g., texture, luminance, structure) on the reliability of diversity and quality metrics?
  • RQ5Are there systematic differences in metric performance when evaluating synthetic images across biomedical versus non-biomedical imaging domains?

Key findings

  • MS-SSIM and CD scores showed no significant variation with sample size for either biomedical or non-biomedical images, indicating metric stability across data quantities.
  • Synthetic X-ray images exhibited better intra-class diversity than real X-ray images, as indicated by higher MS-SSIM and CD scores.
  • Synthetic OCT images showed poorer intra-class diversity than real OCT images, with lower MS-SSIM and CD scores, suggesting difficulty in capturing OCT’s complex features.
  • FID scores varied significantly across imaging modalities due to differences in salient image features, with no consistent trend across classes.
  • The distribution of MS-SSIM, CD, and FID scores differed markedly between biomedical (X-ray, OCT) and non-biomedical (Fashion MNIST) images, reflecting modality-specific feature complexity.
  • The architecture of DCGAN was found to be less effective at generating diverse and high-quality synthetic images for OCT compared to X-ray, due to OCT’s higher feature diversity and complexity.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.