Skip to main content
QUICK REVIEW

[논문 리뷰] Adversarial Attacks on Convolutional Neural Networks in Facial Recognition Domain

Yigit Alparslan, Alparslan, Ken|arXiv (Cornell University)|2020. 01. 30.
Adversarial Robustness in Machine Learning참고 문헌 14인용 수 12
한 줄 요약

이 논문은 컨volutional 신경망(CNN)을 사용한 얼굴 인식 시스템에 대한 적대적 공격을 조사하며, 빠른 기울기 부호 방법(FGSM)과 블랙박스 공격 전략을 적용하여 변형을 생성한다. 실험 결과, 높은 수준의 변형이 존재하더라도 얼굴은 인간이 식별할 수 있지만, 분류기는 확신도 최대 84% 감소하고 정확도 오류율은 81.6%에 이를 정도로 심각한 취약성을 드러내며, DNN의 강건성에 대한 심각한 문제를 제기한다.

ABSTRACT

Numerous recent studies have demonstrated how Deep Neural Network (DNN) classifiers can be fooled by adversarial examples, in which an attacker adds perturbations to an original sample, causing the classifier to misclassify the sample. Adversarial attacks that render DNNs vulnerable in real life represent a serious threat in autonomous vehicles, malware filters, or biometric authentication systems. In this paper, we apply Fast Gradient Sign Method to introduce perturbations to a facial image dataset and then test the output on a different classifier that we trained ourselves, to analyze transferability of this method. Next, we craft a variety of different black-box attack algorithms on a facial image dataset assuming minimal adversarial knowledge, to further assess the robustness of DNNs in facial recognition. While experimenting with different image distortion techniques, we focus on modifying single optimal pixels by a large amount, or modifying all pixels by a smaller amount, or combining these two attack approaches. While our single-pixel attacks achieved about a 15% average decrease in classifier confidence level for the actual class, the all-pixel attacks were more successful and achieved up to an 84% average decrease in confidence, along with an 81.6% misclassification rate, in the case of the attack that we tested with the highest levels of perturbation. Even with these high levels of perturbation, the face images remained identifiable to a human. Understanding how these noised and perturbed images baffle the classification algorithms can yield valuable advances in the training of DNNs against defense-aware adversarial attacks, as well as adaptive noise reduction techniques. We hope our research may help to advance the study of adversarial attacks on DNNs and defensive mechanisms to counteract them, particularly in the facial recognition domain.

연구 동기 및 목표

  • 빠른 기울기 부호 방법(FGSM)을 통해 생성된 적대적 예제가 다양한 얼굴 인식 모델 간에 전이 가능한지를 평가하기 위해.
  • 최소한의 사전 지식(블랙박스 공격) 조건에서 딥 네트워크의 강건성을 얼굴 인식 분야에서 평가하기 위해.
  • 분류기 성능 저하에 있어 단일 픽셀 대비 모든 픽셀 변형 전략의 효과를 조사하기 위해.
  • 시각적으로 최소한의 변형이 존재함에도 불구하고 모델의 확신을 극적으로 감소시킬 수 있는지 분석하기 위해.

제안 방법

  • 얼굴 이미지 데이터셋에 대해 빠른 기울기 부호 방법(FGSM)을 적용하여 적대적 예제를 생성하였다.
  • FGSM로 생성된 변형의 전이성을 테스트하기 위해 별도의 분류기를 훈련시켰다.
  • 대상 모델의 아키텍처나 파라미터에 대한 최소한의 지식을 가정하는 블랙박스 공격 알고리즘을 설계하였다.
  • 두 가지 공격 전략을 탐색: 최적의 단일 픽셀을 크게 변형하거나, 모든 픽셀을 소량으로 변형하는 방식.
  • 두 전략을 결합하여 분류기의 확신도 및 오분류율에 미치는 누적 영향을 평가하였다.
  • 인간 평가를 통해 변형된 이미지가 모델 실패에도 불구하고 여전히 인간이 얼굴로 식별할 수 있음을 확인하였다.

실험 결과

연구 질문

  • RQ1FGSM로 생성된 적대적 예제가 다양한 얼굴 인식 모델 간에 얼마나 전이 가능한가?
  • RQ2공격자가 대상 모델에 대해 최소한의 지식을 가진 상태에서 블랙박스 공격이 얼마나 효과적인가?
  • RQ3단일 픽셀 대비 모든 픽셀 변형 전략이 분류기의 확신도 및 정확도에 미치는 비교적 영향은 어떠한가?
  • RQ4적대적 변형이 인간에게는 인지되지 않게 유지되면서도 모델 성능을 극적으로 악화시킬 수 있는가?
  • RQ5단일 픽셀 및 모든 픽셀 공격 전략을 병합했을 때 얼굴 인식에서 모델의 강건성에 어떤 영향을 미치는가?

주요 결과

  • 단일 픽셀 공격은 정답 클래스에 대해 평균 약 15%의 분류기 확신도 감소를 유도하였다.
  • 모든 픽셀 공격은 높은 변형 수준에서 진짜 클래스에 대해 평균 84%의 확신도 감소를 달성하였다.
  • 가장 효과적인 모든 픽셀 공격은 81.6%의 오분류율을 초래하였으며, 공격 성공도가 높음을 시사하였다.
  • 높은 수준의 변형이 존재하더라도, 모든 적대적으로 변형된 얼굴은 인간 관찰자에게 여전히 식별 가능하였다.
  • 동일한 데이터셋에서 훈련된 다양한 분류기 간에 FGSM로 생성된 적대적 예제의 전이성이 확인되었다.
  • 단일 픽셀 및 모든 픽셀 전략을 병합함으로써 공격의 효과성이 향상되었으며, 이는 모델 성능 저하에 상호보완적인 영향을 미쳤다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.