[论文解读] Analyzing ImageNet with Spectral Relevance Analysis: Towards ImageNet un-Hans'ed.
本文提出了一种可扩展的框架,利用谱相关性分析(SpRAy)检测并缓解 ImageNet 中的 Clever Hans 行为,通过新型基于 Wasserstein 的归因比较方法识别虚假相关性,量化有偏类别,并实施系统性的数据清洗流程,显著降低由人工特征驱动的模型预测。
Today's machine learning models for computer vision are typically trained on very large (benchmark) data sets with millions of samples. These may, however, contain biases, artifacts, or errors that have gone unnoticed and are exploited by the model. In the worst case, the trained model may become a 'Clever Hans' predictor that does not learn a valid and generalizable strategy to solve the problem it was trained for, but bases its decisions on spurious correlations in the training data. Recently developed techniques allow to explain individual model decisions and thus to gain deeper insights into the model's prediction strategies. In this paper, we contribute by providing a comprehensive analysis framework based on a scalable statistical analysis of attributions from explanation methods for large data corpora, here ImageNet. Based on a recent technique - Spectral Relevance Analysis (SpRAy) - we propose three technical contributions and resulting findings: (a) novel similarity metrics based on Wasserstein for comparing attributions to allow for the first time scale, translational, and rotational invariant comparisons of attributions, (b) a scalable quantification of artifactual and poisoned classes where the ML models under study exhibit Clever Hans behavior, (c) a cleaning procedure that allows to relief data of artifacts and biases in a systematic manner yielding significantly reduced Clever Hans behavior, i.e. we un-Hans the ImageNet data corpus. Using this novel method set, we provide qualitative and quantitative analyses of the biases and artifacts in ImageNet and demonstrate that the usage of these insights can give rise to improved models and functionally cleaned data corpora.
研究动机与目标
- 识别并量化大规模数据集(如 ImageNet)中的偏差和人工特征,这些因素导致模型依赖虚假相关性而非学习可泛化的特征。
- 开发一种可扩展的方法,对数百万张图像的模型归因图进行分析,以检测表明 Clever Hans 行为的系统性偏差。
- 通过识别并移除误导模型学习的有缺陷或被污染的类别,实现图像数据集的系统性清洗。
- 通过基于归因分析指导的数据整理,提升模型的鲁棒性和泛化能力,实现对 ImageNet 的‘去 Hans 化’。
提出的方法
- 提出新颖的基于 Wasserstein 的相似性度量,实现对不同图像变换下归因图的尺度、平移和旋转不变比较。
- 应用谱相关性分析(SpRAy)对整个 ImageNet 数据集中的归因图进行大规模计算与分析。
- 引入量化分析流程,基于归因模式识别出具有高水平人工特征或污染信号的类别。
- 开发一种数据清洗流程,通过归因分析指导,移除表现出强烈虚假相关性的图像或类别。
- 利用归因分布的统计建模,检测并隔离模型依赖非语义线索(如背景图案或纹理)的数据区域。
- 通过清洗后模型评估中 Clever Hans 行为的减少程度,验证清洗流程的有效性。
实验结果
研究问题
- RQ1ImageNet 中哪些类别因其训练数据中的虚假相关性而表现出强烈的 Clever Hans 行为证据?
- RQ2如何实现对归因图的比较,使其对尺度、平移和旋转保持不变,从而实现对人工特征的可靠检测?
- RQ3系统性数据清洗在多大程度上能减少在清洗后数据集上微调的模型中的 Clever Hans 行为?
- RQ4对归因图进行谱分析是否能在大规模下揭示 ImageNet 中此前未被发现的偏差或人工特征?
- RQ5移除含人工特征的图像对模型泛化能力和鲁棒性有何影响?
主要发现
- 所提出的基于 Wasserstein 的归因比较度量成功实现了对多样化图像变换下归因图的稳健、不变比较。
- 识别出大量 ImageNet 类别含有高水平的人工特征或污染信号,表明存在广泛存在的 Clever Hans 行为。
- 数据清洗流程有效降低了下游模型中的 Clever Hans 行为,提升了在分布外数据和干净数据上的泛化能力。
- 清洗后模型对非语义线索(如背景图案或纹理)的依赖性降低,证实了虚假相关性的有效缓解。
- 该框架首次实现了基于归因分析的 ImageNet 全局人工特征普遍性的大规模、定量评估。
- 本研究证明,基于可解释性方法指导的系统性数据整理,可生成功能更清洁的数据集,显著提升模型性能与可靠性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。