[论文解读] Detect, Quantify, and Incorporate Dataset Bias: A Neuroimaging Analysis on 12,207 Individuals
本研究通过分析15个数据集中共12,207例T1加权MRI扫描,检测、量化并利用神经影像数据集偏差。研究引入两种度量指标——巴氏距离和年龄预测误差,以衡量数据集兼容性,创建了神经影像采集站点的t-SNE嵌入,揭示了跨数据集的相似性,并证明基于偏差认知的训练集选择可提升自闭症预测的准确性,优于随机采样。
Neuroimaging datasets keep growing in size to address increasingly complex medical questions. However, even the largest datasets today alone are too small for training complex models or for finding genome wide associations. A solution is to grow the sample size by merging data across several datasets. However, bias in datasets complicates this approach and includes additional sources of variation in the data instead. In this work, we combine 15 large neuroimaging datasets to study bias. First, we detect bias by demonstrating that scans can be correctly assigned to a dataset with 73.3% accuracy. Next, we introduce metrics to quantify the compatibility across datasets and to create embeddings of neuroimaging sites. Finally, we incorporate the presence of bias for the selection of a training set for predicting autism. For the quantification of the dataset bias, we introduce two metrics: the Bhattacharyya distance between datasets and the age prediction error. The presented embedding of neuroimaging sites provides an interesting new visualization about the similarity of different sites. This could be used to guide the merging of data sources, while limiting the introduction of unwanted variation. Finally, we demonstrate a clear performance increase when incorporating dataset bias for training set selection in autism prediction. Overall, we believe that the growing amount of neuroimaging data necessitates to incorporate data-driven methods for quantifying dataset bias in future analyses.
研究动机与目标
- 检测并量化大规模神经影像数据集中的数据集偏差,以解决数据融合与模型泛化性受阻的问题。
- 开发基于数据的度量指标,用于衡量神经影像数据集与采集站点之间的兼容性。
- 创建站点级别的嵌入表示,可视化跨数据集的相似性,揭示来自不同数据集的站点可能比同一数据集内的站点更为相似。
- 证明在临床预测任务(如自闭症检测)中,将数据集偏差纳入训练集选择策略可带来实际性能提升。
提出的方法
- 通过训练分类器以预测MRI扫描的来源数据集,实现73.3%的准确率,从而检测数据集偏差。
- 利用脑体积与皮层厚度特征分布之间的巴氏距离,量化数据集兼容性。
- 使用年龄预测误差作为代理指标,量化模型泛化能力,以衡量数据集兼容性。
- 基于成对年龄预测误差,利用t-SNE方法创建神经影像采集站点的嵌入表示,可视化跨数据集的站点相似性。
- 利用站点兼容性度量指标指导非均匀的训练集选择,优先选择度量值较低的站点。
- 基于该度量指标采用指数加权方法,优先在训练集构成中纳入相似站点。
实验结果
研究问题
- RQ1能否基于图像特征可靠地区分不同神经影像数据集,从而表明存在数据集偏差?
- RQ2如何利用基于图像和基于人口统计的度量指标,对数据集兼容性进行定量衡量?
- RQ3来自不同数据集的神经影像站点在图像特征上有多大的相似性?这种相似性能否被有意义地可视化?
- RQ4与随机采样相比,将数据集偏差纳入训练集选择是否能提升自闭症预测的性能?
- RQ5在偏差感知的模型训练中,巴氏距离与年龄预测误差哪个度量指标能带来更好的性能?
主要发现
- 分类器可将MRI扫描准确分配至其来源数据集,准确率达73.3%,证实了显著的数据集偏差存在。
- 在指导自闭症预测的训练集选择方面,年龄预测误差度量指标优于巴氏距离。
- t-SNE可视化揭示了神经影像站点的四个显著聚类,其中来自不同数据集的站点往往比同一数据集内的站点更为相似。
- ABIDE、HCP、GSP和CORR的站点形成一个以年轻受试者为主的聚类,而ADNI和AIBL的站点则形成一个以年长受试者为主的聚类。
- 基于年龄预测误差度量指标的训练集选择,其自闭症分类准确率高于随机采样或基于巴氏距离的选样策略。
- 所提出的偏差感知训练策略显著提升了模型性能,证明了在大规模神经影像分析中量化并整合数据集偏差具有实际效用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。