Skip to main content
QUICK REVIEW

[논문 리뷰] Powerpropagation: A sparsity inducing weight reparameterisation

Jonathan Schwarz, Siddhant M. Jayakumar|arXiv (Cornell University)|2021. 10. 01.
Domain Adaptation and Few-Shot Learning참고 문헌 88인용 수 7
한 줄 요약

Powerpropagation는 기울기 업데이트를 파rameter 크기에 비례하도록 만들어, '부자인 자가 더 많이 얻는' 역동성을 유도함으로써 모델의 희박성(스퍼스리티)을 유도하는 새로운 가중치 재파arameter화 방법이다. 이는 큰 가중치가 더 빠르게 적응하고 작은 가중치는 거의 변화하지 않도록 하여, 0에 가까운 값이 더 많이 분포하는 모델을 생성한다. 이는 더 효과적인 프루닝과 희박 학습 및 계속 학습 설정에서의 성능 향상에 기여한다.

ABSTRACT

The training of sparse neural networks is becoming an increasingly important tool for reducing the computational footprint of models at training and evaluation, as well enabling the effective scaling up of models. Whereas much work over the years has been dedicated to specialised pruning techniques, little attention has been paid to the inherent effect of gradient based training on model sparsity. In this work, we introduce Powerpropagation, a new weight-parameterisation for neural networks that leads to inherently sparse models. Exploiting the behaviour of gradient descent, our method gives rise to weight updates exhibiting a "rich get richer" dynamic, leaving low-magnitude parameters largely unaffected by learning. Models trained in this manner exhibit similar performance, but have a distribution with markedly higher density at zero, allowing more parameters to be pruned safely. Powerpropagation is general, intuitive, cheap and straight-forward to implement and can readily be combined with various other techniques. To highlight its versatility, we explore it in two very different settings: Firstly, following a recent line of work, we investigate its effect on sparse training for resource-constrained settings. Here, we combine Powerpropagation with a traditional weight-pruning technique as well as recent state-of-the-art sparse-to-sparse algorithms, showing superior performance on the ImageNet benchmark. Secondly, we advocate the use of sparsity in overcoming catastrophic forgetting, where compressed representations allow accommodating a large number of tasks at fixed model capacity. In all cases our reparameterisation considerably increases the efficacy of the off-the-shelf methods.

연구 동기 및 목표

  • 신경망의 표준 기울기 기반 학습에서 내재된 희박성 유도의 부족을 해결하기 위해.
  • 학습 과정에서 내재된 희박성을 통해 딥 러닝 모델의 계산 및 메모리 오버헤드를 줄이기 위해.
  • 자연스럽게 희박한 해를 선호하는 파라미터화를 활용하여 기존의 프루닝 및 계속 학습 기법의 효율성을 향상시키기 위해.
  • 자원 제약이 있는 환경에서 더 효율적인 모델 압축을 가능하게 하고, 계속 학습에서 치명적인 잊음 현상을 완화하기 위해.
  • 기존의 희박성 및 압축 기법과 쉽게 통합할 수 있는 단순하고 일반적인 방법을 제공하기 위해.

제안 방법

  • 전방 전파 중에 가중치를 α제곱(α > 1)으로 재파라미터화하여 부호를 유지한다.
  • 이 재파라미터화는 기울기 계산에 |w|^{α-1}의 스케일링 인자를 도입하여, 업데이트가 가중치 크기에 비례하도록 한다.
  • 작은 크기의 가중치는 더 작은 기울기 업데이트를 받으며, 이는 큰 가중치가 학습을 지배하는 '부자인 자가 더 많이 얻는' 역동성을 유도한다.
  • 이 방법은 손실 함수를 수정하거나 추가적인 페널티를 추가하지 않더라도, 희박성을 유도하는 암묵적 정규화 역할을 한다.
  • 표준 최적화와 호환되며, 코드 변경이 최소한으로 필요한 구현 가능하다 (예: 전방 전파에 단 한 줄 추가).
  • 기존 기법들과의 조합이 가능하며, 예를 들어 가중치 프루닝, 스퍼스-투-스퍼스 학습, EfficientPackNet과 같은 효율적인 계속 학습 방법과 함께 사용된다.

실험 결과

연구 질문

  • RQ1기울기 기반 학습을 어떻게 수정하면 명시적 정규화 없이도 내재적으로 더 희박한 모델을 생성할 수 있는가?
  • RQ2Powerpropagation는 훈련된 모델의 희박성 분포와 프루닝 효율성에 어떤 영향을 미치는가?
  • RQ3Powerpropagation는 자원 제약이 있는 환경(예: ImageNet)에서 희박 학습에 성능 향상을 줄 수 있는가?
  • RQ4Powerpropagation는 압축된, 희박한 표현을 통해 희박한 표현을 통해 치명적인 잊음 현상을 줄여 계속 학습에서 성능을 향상시킬 수 있는가?
  • RQ5하이퍼파rameter α는 희박성과 모델 성능 사이의 트레이드오프에 어떤 영향을 미치는가?

주요 결과

  • Powerpropagation로 훈련된 모델는 0에 가까운 가중치의 밀도가 현저히 높아, 더 공격적이고 안전한 프루닝이 가능하다.
  • ImageNet에서 Powerpropagation와 스퍼스-투-스퍼스 학습을 조합한 결과, 동일한 희박성 제약 조건 하에서 기존 방법들을 능가하는 최고 성능을 기록했다.
  • 10개 태스크로 구성된 Split-CIFAR100 벤치마크에서 계속 학습을 수행한 결과, Powerpropagation + EfficientPackNet는 태스크 인fer런스 시 73.70%의 정확도를 달성하여 새로운 최고 기록을 수립했다.
  • 계속되는 세계 강화 학습 벤치마크에서 Powerpropagation는 평균 86%의 성공률을 기록했으며, 거의 0에 가까운 잊음(0.00)을 보이며 이전 방법들을 능가했다.
  • Powerpropagation와 EfficientPackNet의 조합은 Split-MNIST에서 99.71%의 정확도를 달성했으며, 별도의 태스크 헤드를 가진 방법들조차도 능가했다.
  • 이 방법은 희박 학습과 계속 학습을 포함한 다양한 환경에서 매우 효과적이며, 최소한의 구현 오버헤드로도 높은 성능 향상을 이끌어냈다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.