[论文解读] Taking a machine's perspective: Human deciphering of adversarial images.
本研究调查人类是否能够解码对抗性图像——即欺骗人工智能模型的图像——通过测试人们是否能在感知这些图像为无意义内容的情况下,仍能可靠地预测机器的分类结果。在五个对抗性数据集的八个实验中,人类在面对看似无法辨认的图像时,仍能持续识别出模型预测的标签,而非干扰选项,表明人类直觉与机器决策在对抗性情境下存在强烈对齐。
How similar is the human mind to the sophisticated machine-learning systems that mirror its performance? Models of object categorization based on convolutional neural networks (CNNs) have achieved human-level benchmarks in assigning known labels to novel images. These advances promise to support transformative technologies such as autonomous vehicles and machine diagnosis; beyond this, they also serve as candidate models for the visual system itself -- not only in their output but perhaps even in their underlying mechanisms and principles. However, unlike human vision, CNNs can be fooled by adversarial examples -- carefully crafted images that appear as nonsense patterns to humans but are recognized as familiar objects by machines, or that appear as one object to humans and a different object to machines. This seemingly extreme divergence between human and machine classification challenges the promise of these new advances, both as applied image-recognition systems and also as models of the human mind. Surprisingly, however, little work has empirically investigated human classification of such adversarial stimuli: Does human and machine performance fundamentally diverge? Or could humans decipher such images and predict the machine's preferred labels? Here, we show that human and machine classification of adversarial stimuli are robustly related: In eight experiments on five prominent and diverse adversarial imagesets, human subjects reliably identified the machine's chosen label over relevant foils. This pattern persisted for images with strong antecedent identities, and even for images described as totally unrecognizable to human eyes. We suggest that human intuition may be a more reliable guide to machine (mis)classification than has typically been imagined, and we explore the consequences of this result for minds and machines alike.
研究动机与目标
- 调查人类感知是否与对抗性图像上的机器分类一致,挑战此类图像对人类理解完全不透明的假设。
- 测试人类在面对看似为随机噪声或无法辨认的图案的对抗性刺激时,是否能直觉识别出机器选择的标签。
- 探讨这种人类与机器对齐对视觉认知模型以及AI系统在现实应用中可靠性的启示。
- 确定人类直觉在对抗性场景下是否可作为机器行为的可靠预测指标,即使刺激在视觉上不连贯。
提出的方法
- 在包括FGSM、PGD和AutoAttack扰动在内的五个不同对抗性图像数据集上,开展了八个受控实验。
- 向人类参与者展示对抗性图像,并提供多项选择选项,包括模型预测的标签和合理的干扰项。
- 采用强制选择识别任务,评估参与者是否能以高于随机水平的准确率识别出模型的标签。
- 从多个受试者群体收集数据,以确保在不同人群和图像类型中的可推广性。
- 分析在不同感知连贯性和对抗性强度水平下的图像集表现。
- 采用统计建模方法,评估人类表现相对于随机水平的一致性和显著性。
实验结果
研究问题
- RQ1人类能否可靠地识别出对人类观察者而言看似为随机噪声的对抗性图像的机器预测标签?
- RQ2人类在识别模型标签方面的表现是否与对抗性扰动的强度或类型相关?
- RQ3即使图像无法辨认,人类直觉与机器分类之间是否存在系统性关联?
- RQ4人类感知在多大程度上可作为对抗性视觉任务中机器决策的代理?
主要发现
- 在所有八个实验中,人类参与者识别对抗性图像上机器选定标签的表现显著优于随机水平。
- 即使图像被描述为完全无法辨认或视觉上不连贯,人类的准确率依然稳健,表明对细微统计模式具有敏感性。
- 人类与机器分类之间的相关性在不同对抗性图像集和扰动类型中始终保持强烈。
- 参与者即使在图像看似为随机噪声时,也能以高度自信识别出模型的标签,表明对对抗性结构具有感知敏感性。
- 表现不依赖于图像的原始身份(即原始物体类别),表明该效应由对抗性模式本身驱动。
- 结果表明,人类直觉捕捉到了机器决策的关键方面,挑战了人类与机器视觉在对抗性环境下存在根本分歧的观念。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。