Skip to main content
QUICK REVIEW

[논문 리뷰] Exploring the Trade-off between Plausibility, Change Intensity and Adversarial Power in Counterfactual Explanations using Multi-objective Optimization

Javier Del Ser, Alejandro Barredo-Arrieta|arXiv (Cornell University)|2022. 05. 20.
Adversarial Robustness in Machine Learning인용 수 4
한 줄 요약

이 논문은 GAN 기반의 데이터 분포 모델과 다목적 최적화 솔버를 사용하여 타당성, 변화 강도, 적대적 영향력을 균형 잡는 다목적 최적화 프레임워크를 제안한다. 이 프레임워크는 이미지 및 3D 데이터에서 더 신뢰할 수 있고 직관적이며 실행 가능한 반성적 설명을 생성하여 비전문가 사용자를 위한 편향 탐지 및 모델 해석 가능성 향상에 기여한다.

ABSTRACT

There is a broad consensus on the importance of deep learning models in tasks involving complex data. Often, an adequate understanding of these models is required when focusing on the transparency of decisions in human-critical applications. Besides other explainability techniques, trustworthiness can be achieved by using counterfactuals, like the way a human becomes familiar with an unknown process: by understanding the hypothetical circumstances under which the output changes. In this work we argue that automated counterfactual generation should regard several aspects of the produced adversarial instances, not only their adversarial capability. To this end, we present a novel framework for the generation of counterfactual examples which formulates its goal as a multi-objective optimization problem balancing three different objectives: 1) plausibility, i.e., the likeliness of the counterfactual of being possible as per the distribution of the input data; 2) intensity of the changes to the original input; and 3) adversarial power, namely, the variability of the model's output induced by the counterfactual. The framework departs from a target model to be audited and uses a Generative Adversarial Network to model the distribution of input data, together with a multi-objective solver for the discovery of counterfactuals balancing among these objectives. The utility of the framework is showcased over six classification tasks comprising image and three-dimensional data. The experiments verify that the framework unveils counterfactuals that comply with intuition, increasing the trustworthiness of the user, and leading to further insights, such as the detection of bias and data misrepresentation.

연구 동기 및 목표

  • 비전문가 사용자를 대상으로 한 모델 설명 가능성의 격차를 해소하기 위해 직관적이고 실행 가능한 반성적 설명을 생성하기 위해.
  • 타당성, 변화 강도, 적대적 영향력을 균형 잡는 다목적 최적화 문제로 반성적 설명 생성을 체계화하기 위해.
  • 실제 세계 데이터 분포를 반영하고 잠재적 편향을 드러내는 반성적 설명을 생성함으로써 모델의 신뢰성을 향상시키기 위해.
  • 모델 행동에 대한 깊이 있는 통찰을 제공하여 데이터의 잘못된 표현 및 적대적 페르터베이션에 대한 취약성을 파악하기 위해.

제안 방법

  • 입력 데이터 분포를 모델링하기 위해 GAN을 훈련시어, 타당적인 반성적 설명 생성을 가능하게 한다.
  • GAN의 구분자에 속성 벡터를 조건으로 적용하여 원하는 출력 변화를 갖는 반성적 설명의 생성을 유도한다.
  • 세 가지 목표인 타당성, 변화 강도, 적대적 영향력을 균형 잡는 반성적 설명을 찾기 위해 다목적 최적화 솔버를 활용한다.
  • 타당성은 원래 데이터 분포에 가까운 현실적인 샘플 생성 능력으로 평가된다.
  • 변화 강도는 원본 입력과 반성적 설명 간의 L2 거리로 측정된다.
  • 적대적 영향력은 반성적 설명을 타겟 모델에 입력했을 때 출력 변화의 크기로 정량화된다.

실험 결과

연구 질문

  • RQ1반성적 설명 생성은 상충되는 목표를 수반하는 본질적으로 다목적 최적화 문제인가?
  • RQ2생성된 반성적 설명이 데이터 및 작업 맥락에 기반한 직관적 기대에 부합하는가?
  • RQ3다중 기준 반성적 설명은 훈련 세트 내 숨겨진 편향이나 데이터 오류 표현을 드러내는가?
  • RQ4타당성, 변화 강도, 적대적 영향력 간의 트레이드오프가 사용자 신뢰도 및 해석 가능성에 어떤 영향을 미치는가?

주요 결과

  • 여섯 개의 실험에서의 파레토 최적 해 근사치는 반성적 설명 생성이 서로 상충되는 다수의 목표에 의해 지배되며, 이를 해결하기 위해 다목적 최적화가 필요하다는 것을 확인한다.
  • 생성된 반성적 설명은 색상 변화나 구조적 강조와 같은 타당한 변화를 보이며, 인간의 직관과 작업 맥락의 의미론과 일치한다.
  • 이 프레임워크는 구성적 편향과 데이터 오류 표현을 성공적으로 드러내었으며, 특히 속성-클래스 관계에서 잠재적 취약성을 보여주었다.
  • 높은 적대적 영향력과 낮은 변화 강도를 갖는 반성적 설명은 더 실행 가능했고, 모든 데이터셋에서 높은 타당성이 유지되었다.
  • 이 프레임워크는 비전문가 사용자가 모델의 행동을 더 잘 이해할 수 있도록 직관적이고 실행 가능한 예측 대체 방안을 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.