Skip to main content
QUICK REVIEW

[논문 리뷰] Adversarial attacks and defenses in explainable artificial intelligence: A survey

Hubert Baniecki, Przemysław Biecek|arXiv (Cornell University)|2023. 06. 06.
Adversarial Robustness in Machine Learning인용 수 4
한 줄 요약

이 종합 검토는 해석 가능한 인공지능(XAI) 방법과 그 방어 기법에 대한 적대적 공격에 대해 종합적인 분석을 제공하며, 적대적 XAI(AdvXAI) 분야의 연구를 체계화하기 위해 통합된 분류 체계와 기호 체계를 도입한다. 본 논문은 설명 방법의 핵심 취약점을 특정하고, 방어 전략을 평가하며, 고위험 응용 분야에서 XAI의 강건성과 신뢰성 향상을 위한 향후 연구 방향을 제시한다.

ABSTRACT

Explainable artificial intelligence (XAI) methods are portrayed as a remedy for debugging and trusting statistical and deep learning models, as well as interpreting their predictions. However, recent advances in adversarial machine learning (AdvML) highlight the limitations and vulnerabilities of state-of-the-art explanation methods, putting their security and trustworthiness into question. The possibility of manipulating, fooling or fairwashing evidence of the model's reasoning has detrimental consequences when applied in high-stakes decision-making and knowledge discovery. This survey provides a comprehensive overview of research concerning adversarial attacks on explanations of machine learning models, as well as fairness metrics. We introduce a unified notation and taxonomy of methods facilitating a common ground for researchers and practitioners from the intersecting research fields of AdvML and XAI. We discuss how to defend against attacks and design robust interpretation methods. We contribute a list of existing insecurities in XAI and outline the emerging research directions in adversarial XAI (AdvXAI). Future work should address improving explanation methods and evaluation protocols to take into account the reported safety issues.

연구 동기 및 목표

  • 해석 가능한 인공지능(XAI) 방법과 공정성 메트릭스를 겨냥한 적대적 공격에 대한 증가하는 연구를 체계화하기 위해.
  • 특히 고위험 의사결정 환경에서 데이터 및 모델 조작에 기인한 XAI의 주요 실패 모드를 특정하고 분류하기 위해.
  • 적대적 기계학습(AdvML) 및 XAI 연구 공동체 간 격차를 메우기 위해 적대적 XAI에 대한 통합된 기호 체계와 분류 체계를 제안하기 위해.
  • 기존 방어 기법, 특히 모델 정규화, 집중 샘플링, 앙상블 기반 설명 집계를 평가하기 위해.
  • AdvXAI 분야의 새로운 과제, 예를 들어 XAI-세탁, 잘못된 설명에 대한 인간의 취약성, 규제적 영향 등을 부각시키기 위해.

제안 방법

  • 주요 머신러닝 컨ferences(ICML, ICLR, NeurIPS, AAAI, AIj, NMI) 및 그 인용 네트워크에서 50편 이상의 논문을 체계적으로 검토하여 적대적 XAI 분야의 핵심 논문을 특정한다.
  • XAI에 대한 적대적 공격를 위한 통합된 분류 체계와 기호 체계를 도입한다. 이는 데이터 수준, 모델 수준, 설명 수준의 공격를 구분한다.
  • 공격를 그 대상(예: 특성 중요도, 반대사례, 시각화 지도)과 액세스 수준(화이트박스,_BLK박스, GRAY박스)에 따라 분류한다.
  • 방어 전략을 분석하며, 모델 정규화, 데이터 필터링, 단일 방법 조작에 대한 취약성을 줄이는 앙상블 기반 설명 방법을 포함한다.
  • 수정된 설명 방법과 훈련 시점 방어를 통해 모델의 강건성을 향상시키는 전략을 비교 분석한다.
  • 실제 구현 환경에서 정보 과부하 및 속임수 설명에 대한 인간의 취약성과 같은 인간 요소의 영향을 평가한다.
Figure 1: Explainable AI methods have vurlnerabilities related to safety and security, which we call failure modes . They mainly exist on two levels exploiting manipulation of data or the model. Adversarial manipulation of explanations leads to misinterpretation of the model’s behaviour in audits, d
Figure 1: Explainable AI methods have vurlnerabilities related to safety and security, which we call failure modes . They mainly exist on two levels exploiting manipulation of data or the model. Adversarial manipulation of explanations leads to misinterpretation of the model’s behaviour in audits, d

실험 결과

연구 질문

  • RQ1적대적 공격는 사법 결정이나 의료 진단과 같은 고위험 분야에서 XAI 방법의 신뢰성과 보안을 어떻게 위협하는가?
  • RQ2데이터나 모델 파ameter 조작에 노출되었을 때 XAI의 주요 실패 모드는 무엇인가?
  • RQ3모델 정규화, 데이터 필터링, 또는 설명 앙상블과 같은 현재의 방어 기법은 설명에 대한 적대적 위협을 얼마나 효과적으로 완화하는가?
  • RQ4공정성 메트릭스(예: 인구 통합, 동등한 기회)에 대한 공격와 설명 방법에 대한 공격가 교차하는 방식은 무엇이며, 윤리적 AI에 대한 영향은 무엇인가?
  • RQ5특히 인증, 실제 사고 추적, 규제 준수를 고려할 때, 적대적 XAI 분야의 주요 연구 격차와 향후 연구 방향은 무엇인가?

주요 결과

  • 적대적 공격는 모델 예측를 변경하지 않더라도 설명 출력(예: 시각화 지도, 특성 중요도)을 조작할 수 있으며, 이는 모델 설명에 대한 신뢰를 약화시킨다.
  • 모델에 대한 액세스 없이도 블랙박스 공격는 설명 알고리즘의 구조를 악용하는 최적화 기법을 통해 가능하다.
  • 앙상블 기반 설명 방법은 공격자가 일반적으로 한 가지 설명 방법을 겨냥하기 때문에, 적대적 흐트러짐에 대해 더 강건한 편이다.
  • 모델 정규화와 집중된 데이터 샘플링은 모델과 그 설명의 강건성을 향상시키는 효과적인 방어 전략이다.
  • ‘XAI-세탁’의 위험은 점점 증가하고 있으며, 이는 조직이 규제나 법적 요건을 충족시키기 위해 진정성 있고 신뢰할 수 있는 설명을 사용한다고 잘못 주장하는 것을 의미한다.
  • 인간 이해관계자들은 잘못된 또는 과도한 설명에 의해 조작될 수 있으며, 이는 XAI의 강건성에 대한 인간 중심 평가의 필요성을 강조한다.
Figure 2: Adversarial example is the most common attack on local explanations of the image classifier’s prediction. Left [adapted from 40 ] : An original image is classified as a “dog” and its explanation points out to features influencing this decision. The image can be adversarially changed with p
Figure 2: Adversarial example is the most common attack on local explanations of the image classifier’s prediction. Left [adapted from 40 ] : An original image is classified as a “dog” and its explanation points out to features influencing this decision. The image can be adversarially changed with p

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.