[论文解读] Adversarial attacks and defenses in explainable artificial intelligence: A survey
本综述对可解释人工智能(XAI)方法的对抗性攻击及其防御进行了全面分析,提出了一套统一的分类法与符号系统,以系统化对抗性XAI(AdvXAI)研究。它识别出解释方法中的关键脆弱性,评估了防御策略,并提出了未来研究方向,以增强高风险应用场景中XAI的鲁棒性与可信度。
Explainable artificial intelligence (XAI) methods are portrayed as a remedy for debugging and trusting statistical and deep learning models, as well as interpreting their predictions. However, recent advances in adversarial machine learning (AdvML) highlight the limitations and vulnerabilities of state-of-the-art explanation methods, putting their security and trustworthiness into question. The possibility of manipulating, fooling or fairwashing evidence of the model's reasoning has detrimental consequences when applied in high-stakes decision-making and knowledge discovery. This survey provides a comprehensive overview of research concerning adversarial attacks on explanations of machine learning models, as well as fairness metrics. We introduce a unified notation and taxonomy of methods facilitating a common ground for researchers and practitioners from the intersecting research fields of AdvML and XAI. We discuss how to defend against attacks and design robust interpretation methods. We contribute a list of existing insecurities in XAI and outline the emerging research directions in adversarial XAI (AdvXAI). Future work should address improving explanation methods and evaluation protocols to take into account the reported safety issues.
研究动机与目标
- 系统化日益增长的关于针对可解释人工智能(XAI)方法及其公平性度量的对抗性攻击的研究成果。
- 识别并分类在高风险决策场景中,由于数据与模型操纵而引发的XAI关键失效模式。
- 提出统一的符号系统与分类法,弥合对抗性机器学习(AdvML)与XAI研究社群之间的鸿沟。
- 评估现有防御机制,包括模型正则化、聚焦采样以及基于集成的解释聚合方法。
- 突出AdvXAI中的新兴挑战,包括XAI洗白(XAI-washing)、人类对误导性解释的易感性,以及监管影响。
提出的方法
- 本综述对来自主要机器学习会议(ICML、ICLR、NeurIPS、AAAI、AIj、NMI)及其引用网络的50余篇核心论文进行了系统性文献回顾,以识别对抗性XAI领域的核心研究。
- 提出了一套统一的对抗性攻击分类法与符号系统,用于XAI,区分数据级、模型级与解释级攻击。
- 根据攻击目标(如特征重要性、反事实解释、显著性图)与访问级别(白盒、黑盒、灰盒)对攻击进行分类。
- 分析防御策略,包括模型正则化、数据过滤以及通过集成解释方法降低对单一方法操纵的敏感性。
- 比较修补解释方法与通过训练阶段防御提升模型鲁棒性的有效性。
- 评估人为因素(如信息过载、对欺骗性解释的易感性)在实际部署环境中的影响。

实验结果
研究问题
- RQ1对抗性攻击如何在司法裁决或医疗诊断等高风险领域中损害XAI方法的可信度与安全性?
- RQ2当数据或模型参数受到对抗性操纵时,XAI的主要失效模式是什么?
- RQ3当前防御机制(如模型正则化、数据过滤或解释集成)在缓解对解释的对抗性威胁方面有多有效?
- RQ4对公平性度量(如人口均等性、机会均等性)的攻击与对解释方法的攻击如何相互交叉,对伦理AI有何影响?
- RQ5对抗性XAI中的关键研究空白与未来方向是什么,特别是在认证、真实事件追踪与监管合规方面?
主要发现
- 对抗性攻击可在不改变模型预测的前提下操纵XAI输出(如显著性图或特征重要性),从而破坏对模型解释的信任。
- 即使在无模型访问权限的情况下,对解释方法的黑盒攻击也是可行的,攻击者可利用优化技术来利用解释算法的结构特征。
- 基于集成的解释方法对对抗性扰动表现出更强的鲁棒性,因为攻击者通常仅针对单一解释方法。
- 模型正则化与聚焦数据采样是有效的防御策略,可同时提升模型与解释的鲁棒性。
- 存在日益增长的‘XAI洗白’风险,即组织虚假宣称其使用了真实、可信的解释以满足监管或法律要求。
- 人类利益相关者易受误导性或过度解释的影响,凸显了开展以人为中心的XAI鲁棒性评估的必要性。
![Figure 2: Adversarial example is the most common attack on local explanations of the image classifier’s prediction. Left [adapted from 40 ] : An original image is classified as a “dog” and its explanation points out to features influencing this decision. The image can be adversarially changed with p](https://ar5iv.labs.arxiv.org/html/2306.06123/assets/x2.png)
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。