[Paper Review] Evaluating U-net Brain Extraction for Multi-site and Longitudinal Preclinical Stroke Imaging
This study evaluates a U-net convolutional neural network for automated mouse brain extraction across multi-site, longitudinal preclinical stroke MRI data. Trained on 240 multimodal MRI datasets (T2, ADC maps) from six sites and two time points, the U-net achieved 95–97% accuracy with robust generalization across diverse scanners, contrasts, and pathology, outperforming rule-based methods and enabling high-throughput, reliable pipeline integration for preclinical stroke research.
Rodent stroke models are important for evaluating treatments and understanding the pathophysiology and behavioral changes of brain ischemia, and magnetic resonance imaging (MRI) is a valuable tool for measuring outcome in preclinical studies. Brain extraction is an essential first step in most neuroimaging pipelines; however, it can be challenging in the presence of severe pathology and when dataset quality is highly variable. Convolutional neural networks (CNNs) can improve accuracy and reduce operator time, facilitating high throughput preclinical studies. As part of an ongoing preclinical stroke imaging study, we developed a deep-learning mouse brain extraction tool by using a U-net CNN. While previous studies have evaluated U-net architectures, we sought to evaluate their practical performance across data types. We ask how performance is affected with data across: six imaging centers, two time points after experimental stroke, and across four MRI contrasts. We trained, validated, and tested a typical U-net model on 240 multimodal MRI datasets including quantitative multi-echo T2 and apparent diffusivity coefficient (ADC) maps, and performed qualitative evaluation with a large preclinical stroke database (N=1,368). We describe the design and development of this system, and report our findings linking data characteristics to segmentation performance. We consistently found high accuracy and ability of the U-net architecture to generalize performance in a range of 95-97% accuracy, with only modest reductions in performance based on lower fidelity imaging hardware and brain pathology. This work can help inform the design of future preclinical rodent imaging studies and improve their scalability and reliability.
Motivation & Objective
- To develop a robust, automated brain extraction tool for multi-site, longitudinal preclinical rodent stroke MRI data.
- To assess the generalization performance of U-net architectures across varying MRI scanners, contrasts (T2, ADC), and time points post-stroke.
- To overcome limitations of rule-based brain extraction tools that fail on pathological or heterogeneous data.
- To enable high-throughput, scalable, and reproducible neuroimaging pipelines for multi-center preclinical stroke studies.
- To evaluate whether single-channel U-net models can match multi-channel performance, supporting protocol flexibility.
Proposed method
- Trained a U-net CNN on 240 multimodal MRI datasets including quantitative multi-echo T2 and apparent diffusivity coefficient (ADC) maps.
- Used a multi-channel U-net (M-full-multi) and single-channel variants (M-full-single) to evaluate performance across different input configurations.
- Trained models on data from six imaging centers and tested generalization on data from sites not in the training set (e.g., half-site training/testing).
- Applied Dice similarity coefficient (DSC) as the primary metric to quantify segmentation accuracy against manual ground truth.
- Conducted qualitative evaluation on 1,368 scans from the SPAN study, including cases with severe pathology and motion artifacts.
- Performed ablation studies on restricted models trained on subsets of sites or individual contrasts to assess robustness and transferability.
Experimental results
Research questions
- RQ1How does U-net performance vary across different MRI scanners and field strengths in multi-site preclinical stroke imaging?
- RQ2How does brain pathology and morphometric abnormality (e.g., at 2 days post-stroke) affect U-net segmentation accuracy?
- RQ3Can a single U-net model generalize effectively across multiple MRI contrasts (T2, ADC) and time points (early vs. late post-stroke)?
- RQ4How does performance compare between multi-channel and single-channel U-net architectures in the presence of variable data quality?
- RQ5To what extent can a U-net trained on a subset of sites generalize to new, unseen sites in future multi-center studies?
Key findings
- The U-net achieved 95–97% Dice similarity coefficient (DSC) across all tested conditions, with minimal performance degradation despite variable image quality and pathology.
- The multi-channel U-net (M-full-multi) achieved a DSC of 0.957, while the single-channel version (M-full-single) achieved 0.954, showing near-identical performance.
- Performance improved from early (2 days post-stroke) to late (30 days post-stroke) time points, likely due to more regular brain morphology and reduced edema in later stages.
- The model maintained 100% success rate on 206 cases that failed with rule-based brain extraction, demonstrating robustness to severe pathology.
- Restricted models trained on half the sites generalized well to the other half, achieving a mean DSC of 0.953 (multi-channel) and 0.954 (single-channel), indicating strong transferability.
- Models trained on a single contrast (e.g., T2 or ADC) still performed robustly, with the lowest DSC of 0.935 for T2 baseline, confirming cross-contrast generalization.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.