Skip to main content
QUICK REVIEW

[논문 리뷰] Ablation Studies in Artificial Neural Networks

Richard Meyes, Melanie Lu|arXiv (Cornell University)|2019. 01. 24.
Neural Networks and Applications참고 문헌 43인용 수 168
한 줄 요약

이 논문은 MNIST의 얕은 MLP와 ImageNet의 VGG-19 CNN에서 단일 및 쌍 단위 제거를 수행하여 지식 표현, 중복성, 층/클래스 특이적 효과, 손상 후 회복 학습을 연구한다.

ABSTRACT

Ablation studies have been widely used in the field of neuroscience to tackle complex biological systems such as the extensively studied Drosophila central nervous system, the vertebrate brain and more interestingly and most delicately, the human brain. In the past, these kinds of studies were utilized to uncover structure and organization in the brain, i.e. a mapping of features inherent to external stimuli onto different areas of the neocortex. considering the growth in size and complexity of state-of-the-art artificial neural networks (ANNs) and the corresponding growth in complexity of the tasks that are tackled by these networks, the question arises whether ablation studies may be used to investigate these networks for a similar organization of their inner representations. In this paper, we address this question and performed two ablation studies in two fundamentally different ANNs to investigate their inner representations of two well-known benchmark datasets from the computer vision domain. We found that features distinct to the local and global structure of the data are selectively represented in specific parts of the network. Furthermore, some of these representations are redundant, awarding the network a certain robustness to structural damages. We further determined the importance of specific parts of the network for the classification task solely based on the weight structure of single units. Finally, we examined the ability of damaged networks to recover from the consequences of ablations by means of recovery training. We argue that ablations studies are a feasible method to investigate knowledge representations in ANNs and are especially helpful to examine a networks robustness to structural damages, a feature of ANNs that will become increasingly important for future safety-critical applications.

연구 동기 및 목표

  • 현대 인공 신경망에서 지식 표현을 이해하기 위해 신경과학에서 영감을 받은 제거의 사용을 동기화한다.
  • 얕은 MLP와 깊은 CNN에서 제거가 전체 성능 및 클래스별 성능에 어떤 영향을 주는지 검토한다.
  • 네트워크 부분과 층 간 표현의 중복성 및 국지화를 조사한다.
  • 손상 후 회복 학습을 통해 네트워크의 성능 회복 능력을 평가한다.

제안 방법

  • MNIST에서 2-히든 레이어 MLP의 단일 유닛 제거를 각 유닛으로 들어오는 가중치를 0으로 만들어 수행한다.
  • MLP에서 쌍 단위 제거를 수행하여 중복성과 결합 효과를 평가한다.
  • 이미지넷에 대해 사전 학습된 VGG-19의 컨볼루션 필터에 대해 그룹(1%, 5%, 10%, 25%) 제거를 적용하고 유사도 기반 그룹화를 사용한다.
  • ImageNet 검증에서 top-1 및 top-5 정확도로 영향을 평가한다.
  • t-SNE를 사용하여 제거 후 데이터 구조 표현을 시각화한다.
  • 손상 네트워크를 재학습하여 재파손 회복 학습으로 성능 회복을 측정한다.

실험 결과

연구 질문

  • RQ1제거가 신경과학의 발견과 유사하게 ANN에서 국소적 표현 대 분산 표현을 드러내는가?
  • RQ2단일 유닛 및 쌍 제거가 얕은 아키텍처와 깊은 아키텍처의 전체 성능 및 클래스별 성능에 어떤 영향을 미치는가?
  • RQ3일부 층이나 클래스가 제거에 더 취약하거나 더 강인한가?
  • RQ4손상된 네트워크가 타깃 회복 학습을 통해 성능을 회복할 수 있으며 어느 정도까지 가능한가?
  • RQ5제거에 대한 버퍼링이 되는 중복성이 있고 특정 클래스를 위한 선택적 표현이 있는가?

주요 결과

  • MLP의 단일 유닛 제거는 전체 정확도에 큰 하락을 일으킬 수 있다(예: 최대 44.5 퍼센트 포인트). 그러나 일부 클래스의 정확도는 개선될 수 있다.
  • 일부 유닛은 보편적으로 중요하고 다른 유닛은 특정 클래스에 대해 선택적으로 중요한데, 그 영향은 학습 후 들어오는 가중치 분포의 변화와 상관관계가 있다.
  • 쌍 제거는 단일 제거의 합보다 더 큰 효과를 낼 수 있어 중복성이 존재하지만 유닛 간의 잠재적 고유 상호작용도 있음을 시사한다.
  • VGG-19에서 제거 효과는 층에 따라 다르며 더 큰 비율의 제거일수록 정확도 하락이 커지는 경향이 있다; 일부 층은 특정 클래스에 더 중요한 역할을 한다.
  • 일부 층에서 제거가 일부 클래스를 위한 성능을 예기치 않게 증가시켜 표현의 복잡한 재분배를 시사한다.
  • 회복 학습은 제거 후 성능을 크게 회복할 수 있으며, 때로는 심한 손상(예: 한 층의 필터 약 80% 제거) 이후에도 거의 원래 수준으로 회복된다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.