[논문 리뷰] Chain-NN: An Energy-Efficient 1D Chain Architecture for Accelerating Deep Convolutional Neural Networks
Chain-NN는 식각 기반 원리와 열 방향 입력 스캔을 통해 데이터 이동을 최소화함으로써 딥 컨volution 신경망(CNN)의 가속을 위한 이원채널 처리 엔진을 갖춘 1차원 체인 아키텍처를 제안한다. TSMC 28nm 공정에 구현된 576-PE 버전은 567.5mW에서 806.4 GOPS의 성능을 달성하며, 에너지 효율성은 1421.0 GOPS/W로 최신 기술 대비 2.5~4.1배 향상되었다.
Deep convolutional neural networks (CNN) have shown their good performances in many computer vision tasks. However, the high computational complexity of CNN involves a huge amount of data movements between the computational processor core and memory hierarchy which occupies the major of the power consumption. This paper presents Chain-NN, a novel energy-efficient 1D chain architecture for accelerating deep CNNs. Chain-NN consists of the dedicated dual-channel process engines (PE). In Chain-NN, convolutions are done by the 1D systolic primitives composed of a group of adjacent PEs. These systolic primitives, together with the proposed column-wise scan input pattern, can fully reuse input operand to reduce the memory bandwidth requirement for energy saving. Moreover, the 1D chain architecture allows the systolic primitives to be easily reconfigured according to specific CNN parameters with fewer design complexity. The synthesis and layout of Chain-NN is under TSMC 28nm process. It costs 3751k logic gates and 352KB on-chip memory. The results show a 576-PE Chain-NN can be scaled up to 700MHz. This achieves a peak throughput of 806.4GOPS with 567.5mW and is able to accelerate the five convolutional layers in AlexNet at a frame rate of 326.2fps. 1421.0GOPS/W power efficiency is at least 2.5 to 4.1x times better than the state-of-the-art works.
연구 동기 및 목표
- 외부 메모리 대역폭을 최소화하여 딥 CNN 추론의 에너지 소비를 감소시키기 위해.
- CNN 가속기에서 프로세서와 메모리 계층 간의 데이터 이동에 따른 높은 전력 소모 문제를 해결하기 위해.
- 다양한 CNN 레이어 구성에 효율적으로 대응할 수 있는 확장성 있고 재구성 가능한 아키텍처를 설계하기 위해.
- 자원 제약이 있는 플랫폼에서 실시간 CNN 추론을 위한 높은 처리량과 에너지 효율성을 달성하기 위해.
제안 방법
- 컨volution 연산을 위해 이원채널 식각 배열로 구성된 1차원 처리 요소(PE) 체인을 활용한다.
- 연속된 PE들로 구성된 1차원 식각 원리를 사용하여 자원 재사용 최적화 방식으로 컨volution을 수행한다.
- 입력 특징 맵의 재사용을 극대화하고 메모리 액세스를 줄이기 위해 열 단위 스캔 입력 패턴을 도입한다.
- 다양한 CNN 레이어 차원에 쉽게 재구성 가능하도록 설계하여 설계 복잡도를 최소화한다.
- TSMC 28nm 공정에 3751k 논리 게이트와 352KB 온칩 메모리를 구현하여 면적과 전력 사용 효율성을 극대화한다.
실험 결과
연구 질문
- RQ1CNN 가속기에서 데이터 이동을 어떻게 최소화하여 에너지 소비를 줄일 수 있는가?
- RQ2식각 원리를 활용한 1차원 체인 아키텍처가 다양한 CNN 레이어에 대해 높은 처리량과 에너지 효율성을 달성할 수 있는가?
- RQ3열 단위 입력 스캔 방식이 데이터 재사용과 메모리 대역폭 감소에 얼마나 기여하는가?
- RQ41차원 체인 아키텍처의 재구성 가능성은 설계 복잡도와 다양한 네트워크에 대한 적응성에 어떤 영향을 미치는가?
주요 결과
- 576-PE 기반 Chain-NN는 TSMC 28nm 공정에서 700MHz에서 최대 806.4 GOPS의 처리량을 달성한다.
- 설계는 오직 567.5mW의 전력을 소비하여 1421.0 GOPS/W의 전력 효율성을 확보한다.
- 에너지 효율성은 최신 기술 대비 2.5배에서 4.1배 높다.
- 아카이브넷의 다섯 번째 컨볼루션 레이어를 모두 326.2fps로 가속화하여 실시간 추론을 가능하게 한다.
- 열 단위 스캔 입력 패턴은 더 높은 연산자 재사용을 통해 메모리 대역폭 요구량을 크게 감소시킨다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.