Skip to main content
QUICK REVIEW

[논문 리뷰] Towards Memory- and Time-Efficient Backpropagation for Training Spiking Neural Networks

Qingyan Meng, Mingqing Xiao|arXiv (Cornell University)|2023. 02. 28.
Advanced Memory and Neural Computing인용 수 4
한 줄 요약

이 논문은 스파iking 신경망(SNNs)을 훈련시키기 위한 메모리 및 시간 효율적인 역전파 방법인 시간을 통한 공간적 학습(SLTT)을 제안한다. 역전파 과정에서 중요하지 않은 시간적 경로를 선택적으로 무시함으로써 SLTT는 스칼라 곱셈의 수를 줄이고, 시퀀스 길이에 관계없이 일정한 메모리 사용량을 달성한다. ImageNet에서 SLTT는 보조 기울기와 함께 BPTT보다 70퍼센트 이상 낮은 메모리 비용과 50퍼센트 빠른 훈련 속도로 최신 기술 수준의 정확도를 달성한다.

ABSTRACT

Spiking Neural Networks (SNNs) are promising energy-efficient models for neuromorphic computing. For training the non-differentiable SNN models, the backpropagation through time (BPTT) with surrogate gradients (SG) method has achieved high performance. However, this method suffers from considerable memory cost and training time during training. In this paper, we propose the Spatial Learning Through Time (SLTT) method that can achieve high performance while greatly improving training efficiency compared with BPTT. First, we show that the backpropagation of SNNs through the temporal domain contributes just a little to the final calculated gradients. Thus, we propose to ignore the unimportant routes in the computational graph during backpropagation. The proposed method reduces the number of scalar multiplications and achieves a small memory occupation that is independent of the total time steps. Furthermore, we propose a variant of SLTT, called SLTT-K, that allows backpropagation only at K time steps, then the required number of scalar multiplications is further reduced and is independent of the total time steps. Experiments on both static and neuromorphic datasets demonstrate superior training efficiency and performance of our SLTT. In particular, our method achieves state-of-the-art accuracy on ImageNet, while the memory cost and training time are reduced by more than 70% and 50%, respectively, compared with BPTT.

연구 동기 및 목표

  • 스파iking 신경망(SNNs)에서 보조 기울기를 사용한 역전파를 통한 시간(BPTT)의 높은 메모리 및 훈련 시간 비용을 해결하기 위해.
  • 역전파 과정에서 계산 그래프 내 중요하지 않은 시간적 경로를 식별하고 제거하여 스칼라 곱셈의 수를 줄이기 위해.
  • 전체 시퀀스 상태를 저장하는 데 의존하지 않고 온라인, 인크리멘탈 기울기 계산을 가능하게 하여 시퀀스 길이에 관계없이 일정한 메모리 사용량을 달성하기 위해.
  • K개의 선택된 시간 단계에서만 역전파를 수행하는 SLTT-K라는 변형을 개발하여 계산 복잡도를 추가로 낮추되 성능 저하 없이 수행하기 위해.
  • 기본 데이터셋에서 최신 기술 수준의 SNN 성능을 달성하면서도 훈련 효율성을 크게 향상시키기 위해.

제안 방법

  • BPTT에서 유도된 기울기를 공간적 및 시간적 성분으로 분해하여 오차 역전파에 대한 시간적 기여도를 명시적으로 분석한다.
  • 역전파 과정에서 중요하지 않은 시간 경로를 제거하는 프루닝 전략을 도입하여 스칼라 곱셈의 수를 줄인다.
  • 각 시간 단계에서 기울기를 즉각적으로 계산함으로써 온라인 훈련을 가능하게 하여 시간 단계 간 중간 상태를 저장할 필요 없이 처리한다.
  • SLTT-K 변형은 역전파를 오직 K개의 시간 단계로 제한하여 시간 복잡도를 Ω(T)에서 Ω(K)로 낮춘다. 여기서 T는 총 시퀀스 길이다.
  • 시간 단계별 배치 정규화를 사용하고 최대 풀링을 평균 풀링으로 대체하여 온라인 학습을 지원하고 훈련 안정성을 향상시킨다.
  • 표준 SNN 아키텍처인 VGG-11, ResNet, 정규화 없는 ResNet(NF-ResNet)과 호환되며 정적 및 뉴로모픽 데이터셋 모두를 지원한다.
Figure 1 : The training time and memory cost comparison between the proposed SLTT-1 method and the BPTT with SG method on ImageNet. SLTT-1 achieves similar accuracy as BPTT, while owning better training efficiency than BPTT both theoretically and experimentally. Please refer to Sections 4 and 5 for
Figure 1 : The training time and memory cost comparison between the proposed SLTT-1 method and the BPTT with SG method on ImageNet. SLTT-1 achieves similar accuracy as BPTT, while owning better training efficiency than BPTT both theoretically and experimentally. Please refer to Sections 4 and 5 for

실험 결과

연구 질문

  • RQ1SNN의 역전파에서 시간적 의존성이 최종 기울기에 얼마나 기여하는가?
  • RQ2계산 그래프 내 중요하지 않은 시간적 경로를 역전파 과정에서 안전하게 프루닝할 수 있는가, 성능 저하 없이?
  • RQ3전체 시퀀스 상태를 저장하지 않고도 SNN에서 온라인 기울기 계산을 달성할 수 있는가, 이로 인해 시퀀스 길이에 관계없이 일정한 메모리 사용량을 달성할 수 있는가?
  • RQ4역전파를 오직 K개의 시간 단계로 제한할 경우 훈련 효율성과 모델 정확도에 어떤 영향을 미치는가?
  • RQ5제안된 방법이 ImageNet과 같은 대규모 데이터셋에서 최신 기술 수준의 성능을 달성하면서도 메모리 및 시간 비용을 크게 줄일 수 있는가?

주요 결과

  • ImageNet에서 SLTT는 보조 기울기와 함께 BPTT보다 70퍼센트 이상 낮은 메모리 비용과 50퍼센트 이상 빠른 훈련 시간을 기록하면서 최신 기술 수준의 정확도를 달성한다.
  • SLTT-1 변형은 ImageNet, CIFAR-10, CIFAR-100, DVS-Gesture, DVS-CIFAR10에서 BPTT 성능을 그대로 유지하면서 훈련 효율성을 극적으로 향상시킨다.
  • SLTT-K는 스칼라 곱셈의 수를 Ω(T)에서 Ω(K)로 줄여 추가적인 시간 복잡도 감소를 가능하게 하며 성능 저하 없이 수행된다.
  • 정적 및 뉴로모픽 데이터셋 모두에서 경쟁적인 성능을 달성하며, 6단계 시퀀스로 ImageNet에서 최신 기술 수준의 결과를 기록한다.
  • SLTT의 온라인 훈련 프레임워크는 중간 상태를 저장할 필요 없이 시퀀스 길이에 관계없이 일정한 메모리 사용량을 가능하게 한다.
  • 이 방법은 VGG-11, NF-ResNet-34/50/101 등 다양한 SNN 아키텍처와 호환되며, 다양한 데이터셋에 적합한 교차 엔트로피 및 MSE 기반 손실 함수를 모두 지원한다.
Figure 2 : Computational graph of multi-layer SNNs. Dashed arrows represent the non-differentiable spike generation functions.
Figure 2 : Computational graph of multi-layer SNNs. Dashed arrows represent the non-differentiable spike generation functions.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.