Skip to main content
QUICK REVIEW

[Paper Review] Detect, Quantify, and Incorporate Dataset Bias: A Neuroimaging Analysis on 12,207 Individuals

Christian Wachinger, Benjamín Gutiérrez-Becker|arXiv (Cornell University)|Apr 28, 2018
Functional Brain Connectivity Studies25 references4 citations
TL;DR

This study detects, quantifies, and leverages dataset bias in neuroimaging by analyzing 12,207 T1-weighted MRI scans across 15 datasets. It introduces two metrics—Bhattacharyya distance and age prediction error—to measure dataset compatibility, creates a t-SNE embedding of neuroimaging sites revealing cross-dataset similarities, and demonstrates that bias-informed training set selection improves autism prediction accuracy over random sampling.

ABSTRACT

Neuroimaging datasets keep growing in size to address increasingly complex medical questions. However, even the largest datasets today alone are too small for training complex models or for finding genome wide associations. A solution is to grow the sample size by merging data across several datasets. However, bias in datasets complicates this approach and includes additional sources of variation in the data instead. In this work, we combine 15 large neuroimaging datasets to study bias. First, we detect bias by demonstrating that scans can be correctly assigned to a dataset with 73.3% accuracy. Next, we introduce metrics to quantify the compatibility across datasets and to create embeddings of neuroimaging sites. Finally, we incorporate the presence of bias for the selection of a training set for predicting autism. For the quantification of the dataset bias, we introduce two metrics: the Bhattacharyya distance between datasets and the age prediction error. The presented embedding of neuroimaging sites provides an interesting new visualization about the similarity of different sites. This could be used to guide the merging of data sources, while limiting the introduction of unwanted variation. Finally, we demonstrate a clear performance increase when incorporating dataset bias for training set selection in autism prediction. Overall, we believe that the growing amount of neuroimaging data necessitates to incorporate data-driven methods for quantifying dataset bias in future analyses.

Motivation & Objective

  • To detect and quantify dataset bias in large-scale neuroimaging datasets, which hinders data fusion and model generalization.
  • To develop data-driven metrics for measuring compatibility between neuroimaging datasets and acquisition sites.
  • To create a site-level embedding that visualizes similarity across datasets, revealing that sites from different datasets can be more similar than those within the same dataset.
  • To demonstrate the practical benefit of incorporating dataset bias into training set selection for clinical prediction tasks, such as autism detection.

Proposed method

  • Detects dataset bias by training a classifier to predict the source dataset of MRI scans, achieving 73.3% accuracy.
  • Quantifies dataset compatibility using the Bhattacharyya distance between feature distributions of brain volume and cortical thickness.
  • Quantifies model generalization using age prediction error as a proxy for dataset compatibility.
  • Creates a t-SNE-based embedding of neuroimaging sites using pairwise age prediction error to visualize site similarity across datasets.
  • Uses the site compatibility metric to guide non-uniform training set selection for autism classification, favoring sites with lower metric values.
  • Employs exponential weighting based on the metric to prioritize similar sites in training set composition.

Experimental results

Research questions

  • RQ1Can neuroimaging datasets be reliably distinguished from one another based on image features, indicating the presence of dataset bias?
  • RQ2How can dataset compatibility be quantitatively measured using image-based and demographic-based metrics?
  • RQ3To what extent do neuroimaging sites from different datasets resemble one another in terms of image characteristics, and can this be visualized meaningfully?
  • RQ4Does incorporating dataset bias into training set selection improve performance in autism prediction compared to random sampling?
  • RQ5Which metric—Bhattacharyya distance or age prediction error—yields better performance in bias-aware model training?

Key findings

  • A classifier could assign MRI scans to their source dataset with 73.3% accuracy, confirming the presence of significant dataset bias.
  • The age prediction error metric outperformed the Bhattacharyya distance in guiding training set selection for autism prediction.
  • t-SNE visualization revealed four distinct clusters of neuroimaging sites, with sites from different datasets often being more similar than sites within the same dataset.
  • Sites from ABIDE, HCP, GSP, and CORR formed a cluster of younger subjects, while ADNI and AIBL formed a cluster of older subjects.
  • Training set selection based on the age prediction error metric led to higher autism classification accuracy than random sampling or Bhattacharyya distance-based selection.
  • The proposed bias-aware training strategy improved model performance, demonstrating the utility of quantifying and incorporating dataset bias in large-scale neuroimaging analyses.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.