[Paper Review] Machine Learning with Multi-Site Imaging Data: An Empirical Study on the Impact of Scanner Effects
The paper shows that scanner/site differences persist after standard neuroimaging pre-processing and can be exploited by classifiers, highlighting challenges in harmonizing multi-site imaging data for machine learning.
This is an empirical study to investigate the impact of scanner effects when using machine learning on multi-site neuroimaging data. We utilize structural T1-weighted brain MRI obtained from two different studies, Cam-CAN and UK Biobank. For the purpose of our investigation, we construct a dataset consisting of brain scans from 592 age- and sex-matched individuals, 296 subjects from each original study. Our results demonstrate that even after careful pre-processing with state-of-the-art neuroimaging pipelines a classifier can easily distinguish between the origin of the data with very high accuracy. Our analysis on the example application of sex classification suggests that current approaches to harmonize data are unable to remove scanner-specific bias leading to overly optimistic performance estimates and poor generalization. We conclude that multi-site data harmonization remains an open challenge and particular care needs to be taken when using such data with advanced machine learning methods for predictive modelling.
Motivation & Objective
- Demonstrate that multi-site T1-weighted MRI data retain scanner-specific bias after state-of-the-art pre-processing.
- Quantify the ability to classify data origin (site) from processed images and tissue maps.
- Assess how data harmonization approaches affect predictive modeling tasks like sex classification.
Proposed method
- Construct a balanced, age- and sex-matched dataset from Cam-CAN and UK Biobank (n=592, 296 per study).
- Apply a common pre-processing pipeline (reorientation, skull-stripping, bias correction, registration, whitening) and generate tissue probability maps with SPM12 and FAST.
- Train random forest classifiers to distinguish data origin and to perform sex classification under various data arrangements (single-site vs multi-site).
- Evaluate site-predictive power and sex-classification performance with cross-validation and reporting of accuracy, entropy, and predicted probabilities.
Experimental results
Research questions
- RQ1Can scanner/site differences be recovered from pre-processed MRI data and derived tissue maps?
- RQ2To what extent does data harmonization reduce site-specific bias in multi-site MRI datasets?
- RQ3How does multi-site data affect the accuracy and generalization of sex classification tasks?
- RQ4What is the impact of varying alignment/normalization on residual scanner effects?
Key findings
- Site classification succeeds with high accuracy even after careful pre-processing, indicating persistent scanner effects.
- Derived tissue probability maps retain scanner bias, and higher spatial normalization can amplify these effects.
- Multi-site age/sex-mmatched data yields sex-classification accuracy similar to single-site data, but sex imbalance and cross-site testing reveal generalization issues.
- Affine registration removing brain size information can worsen drop in classification performance across sites.
- When mixing sites, some configurations (e.g., Cam-CAN females vs UKBB males) yield very high accuracy, suggesting strong site-specific cues persist.
- Overall, data harmonization for multi-site neuroimaging remains challenging and can lead to optimistic performance estimates if not properly addressed.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.