Skip to main content
QUICK REVIEW

[논문 리뷰] A Survey on Deep Neural Network Compression: Challenges, Overview, and Solutions

Rahul Mishra, Hari Prabhat Gupta|arXiv (Cornell University)|2020. 10. 05.
Advanced Neural Network Applications참고 문헌 97인용 수 60
한 줄 요약

이 논문은 기존 심층 신경망(DNN) 압축 기법들을 조사하고 이를 가지치기, 희소 표현, 저정밀도, 지식 증류, 및 기타로 분류하며, IoT 배치를 위한 도전 과제와 향후 방향에 대해 논의한다.

ABSTRACT

Deep Neural Network (DNN) has gained unprecedented performance due to its automated feature extraction capability. This high order performance leads to significant incorporation of DNN models in different Internet of Things (IoT) applications in the past decade. However, the colossal requirement of computation, energy, and storage of DNN models make their deployment prohibitive on resource constraint IoT devices. Therefore, several compression techniques were proposed in recent years for reducing the storage and computation requirements of the DNN model. These techniques on DNN compression have utilized a different perspective for compressing DNN with minimal accuracy compromise. It encourages us to make a comprehensive overview of the DNN compression techniques. In this paper, we present a comprehensive review of existing literature on compressing DNN model that reduces both storage and computation requirements. We divide the existing approaches into five broad categories, i.e., network pruning, sparse representation, bits precision, knowledge distillation, and miscellaneous, based upon the mechanism incorporated for compressing the DNN model. The paper also discussed the challenges associated with each category of DNN compression techniques. Finally, we provide a quick summary of existing work under each category with the future direction in DNN compression.

연구 동기 및 목표

  • 자원 제약이 큰 IoT 기기에 배포를 가능하게 하기 위한 DNN 압축 기법에 대한 철저한 개요를 제공한다.
  • 압축 방법을 다섯 가지 큰 범주로 분류하고 각 범주 내 대표적인 연구들을 매핑한다.
  • 현재 기법의 도전과제와 격차를 식별하여 향후 연구 방향을 제시한다.
  • 각 범주가 정확도 손실을 최소화하면서 저장 공간, 계산, 에너지 요구를 어떻게 줄이는지 요약한다.

제안 방법

  • DNN 압축 기법을 다섯 가지 범주로 분류한다: 네트워크 가지치기, 희소 표현, 비트 정밀도, 지식 증류, 및 기타.
  • 각 범주 내 하위 범주를 검토한다(예: 채널/필터/연결/레이어 가지치기; 양자화, 다중화, 가중치 공유; 정수 추정, 저비트 표현, 이진화; 로짓 전이, 교사 보조, 도메인 적응).
  • 각 범주와 관련된 도전과제(정확도 트레이드오프 및 자원 제약 기기에서의 배치 고려사항 포함)를 논의한다.
  • 각 범주 아래의 기존 연구를 통합적으로 요약하고 DNN 압축의 향후 방향을 제시한다.

실험 결과

연구 질문

  • RQ1문헌에서 DNN 압축 기법의 주요 범주와 하위 범주는 무엇인가?
  • RQ2각 압축 범주와 관련된 핵심 도전과 정확도 트레이드오프는 무엇인가?
  • RQ3현재 접근 방식은 자원 제약 IoT 기기의 배치를 어떻게 다루고 있으며, 미래 연구를 위한 간극은 어디에 있는가?
  • RQ4현실적인 정확도 손실 없이 저장, 계산, 에너지 효율성을 향상시키기 위해 DNN 압축을 어떻게 더 발전시킬 수 있는가?

주요 결과

  • DNN 압축 문헌은 다섯 가지 큰 범주로 조직된다: 네트워크 가지치기, 희소 표현, 비트 정밀도, 지식 증류, 및 기타.
  • 네트워크 가지치기의 주요 하위 범주는 채널, 필터, 연결, 및 레이어 가지치기로 각각 고유한 전략과 트레이드오프를 가진다.
  • 희소 표현에는 저장 공간과 FLOP를 줄이되 성능을 유지하기 위한 양자화, 다중화, 가중치 공유가 포함된다.
  • 비트 정밀도 기법은 가중치 저장 및 계산을 줄이기 위해 정수 추정, 저비트 표현, 이진화를 다룬다.
  • 지식 증류는 큰 teacher 모델에서 작은 student 모델로 일반화를 전이시켜 압축 후 정확도 손실을 완화한다.
  • 기타 기술은 모바일 및 임베디드 기기 적합성 및 병렬화와 같은 배치 측면에 초점을 맞춘다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.