Skip to main content
QUICK REVIEW

[论文解读] Machine Learning with Multi-Site Imaging Data: An Empirical Study on the Impact of Scanner Effects

Ben Glocker, R. H. Robinson|arXiv (Cornell University)|Oct 10, 2019
Radiomics and Machine Learning in Medical Imaging参考文献 15被引用 70
一句话总结

该论文显示,即使经过标准神经影像预处理,扫描仪/站点差异仍然存在,并且可被分类器利用,凸显在多站点成像数据的机器学习数据 harmonization 面临的挑战。

ABSTRACT

This is an empirical study to investigate the impact of scanner effects when using machine learning on multi-site neuroimaging data. We utilize structural T1-weighted brain MRI obtained from two different studies, Cam-CAN and UK Biobank. For the purpose of our investigation, we construct a dataset consisting of brain scans from 592 age- and sex-matched individuals, 296 subjects from each original study. Our results demonstrate that even after careful pre-processing with state-of-the-art neuroimaging pipelines a classifier can easily distinguish between the origin of the data with very high accuracy. Our analysis on the example application of sex classification suggests that current approaches to harmonize data are unable to remove scanner-specific bias leading to overly optimistic performance estimates and poor generalization. We conclude that multi-site data harmonization remains an open challenge and particular care needs to be taken when using such data with advanced machine learning methods for predictive modelling.

研究动机与目标

  • 证明多站点的 T1 加权 MRI 数据在最先进的预处理后仍保留扫描仪特异性偏差。
  • 量化在处理后的图像和组织概率图中推断数据来源(站点)的能力。
  • 评估数据 harmonization 方法对性别分类等预测建模任务的影响。

提出的方法

  • 从 Cam-CAN 和 UK Biobank 构建一个平衡的、按年龄和性别匹配的数据集(n=592,每项研究 296)。
  • 应用通用预处理流程(重新定向、颅骨剥离、偏差校正、配准、白化)并使用 SPM12 与 FAST 生成组织概率图。
  • 训练随机森林分类器以区分数据来源并在不同数据排列(单站点 vs 多站点)下执行性别分类。
  • 使用交叉验证评估站点预测能力和性别分类性能,并报告准确率、熵值与预测概率。

实验结果

研究问题

  • RQ1预处理后的 MRI 数据及派生的组织概率图是否能恢复出扫描仪/站点差异?
  • RQ2数据 harmonization 在多站点 MRI 数据集中在多大程度上减少站点特异性偏差?
  • RQ3多站点数据如何影响性别分类任务的准确性与泛化能力?
  • RQ4不同对齐/归一化对残留扫描仪效应有何影响?

主要发现

  • 即使在仔细的预处理后,站点分类仍以较高的准确率成功,表明扫描仪效应持续存在。
  • 派生的组织概率图仍保留扫描仪偏差,且更高的空间归一化会放大这些效应。
  • 多站点年龄/性别匹配数据的性别分类准确性与单站点数据相近,但性别不平衡与跨站点测试揭示泛化问题。
  • 进行脑容量信息移除的仿射配准可能在跨站点时进一步降低分类性能。
  • 混合站点时,一些配置(如 Cam-CAN 女性 vs UKBB 男性)显示极高的准确率,提示仍存在强烈的站点特定线索。
  • 总体而言,多站点神经影像数据的 harmonization 仍然具有挑战性,若未妥善处理,可能导致乐观的性能估计。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。