Skip to main content
QUICK REVIEW

[논문 리뷰] A Unified Framework for Soft Threshold Pruning

Yanqi Chen, Zhengyu Ma|arXiv (Cornell University)|2023. 02. 25.
Sparse and Compressive Sensing Techniques인용 수 6
한 줄 요약

이 논문은 소프트 스레시홀드 프루닝을 반복적 수축-임계값 알고리즘(ISTA)과 연결하는 통합 이론적 프레임워크를 제안한다. 이를 통해 L1-정규화 최적화 문제의 암묵적 해로 프루닝을 재구성한다. L1 정규화 계수를 안정화하는 최적의 시간 불변 임계값 스케줄러를 유도함으로써, 하이퍼파ram터 조정을 최소화하고 다양한 SGD로 훈련된 모델에 광범위하게 적용 가능한 상태의 성능을 달성한다. ResNet-50, MobileNet-V1, 그리고 ImageNet 기반 SEW ResNet-18에서 최고의 프루닝 성능을 기록한다.

ABSTRACT

Soft threshold pruning is among the cutting-edge pruning methods with state-of-the-art performance. However, previous methods either perform aimless searching on the threshold scheduler or simply set the threshold trainable, lacking theoretical explanation from a unified perspective. In this work, we reformulate soft threshold pruning as an implicit optimization problem solved using the Iterative Shrinkage-Thresholding Algorithm (ISTA), a classic method from the fields of sparse recovery and compressed sensing. Under this theoretical framework, all threshold tuning strategies proposed in previous studies of soft threshold pruning are concluded as different styles of tuning $L_1$-regularization term. We further derive an optimal threshold scheduler through an in-depth study of threshold scheduling based on our framework. This scheduler keeps $L_1$-regularization coefficient stable, implying a time-invariant objective function from the perspective of optimization. In principle, the derived pruning algorithm could sparsify any mathematical model trained via SGD. We conduct extensive experiments and verify its state-of-the-art performance on both Artificial Neural Networks (ResNet-50 and MobileNet-V1) and Spiking Neural Networks (SEW ResNet-18) on ImageNet datasets. On the basis of this framework, we derive a family of pruning methods, including sparsify-during-training, early pruning, and pruning at initialization. The code is available at https://github.com/Yanqi-Chen/LATS.

연구 동기 및 목표

  • 기존 소프트 스레시홀드 프루닝 방법들이 부호화된 또는 학습 가능한 임계값 스케줄링에 의존함에 따라 이론적 기반의 부족을 해결하기 위해.
  • 다양한 임계값 조정 전략을 하나의 최적화 프레임워크로 통합하기 위해.
  • 시간 불변 목적 함수를 보장하기 위해 L1 정규화 계수를 일정하게 유지하는 최적의 임계값 스케줄러를 유도하기 위해.
  • 모델 특화 조정 없이도 인공신경망과 스파iking 신경망을 포함한 다양한 모델에 체계적인 프루닝을 가능하게 하기 위해.

제안 방법

  • 반복적 수축-임계값 알고리즘(ISTA)을 사용해 소프트 스레시홀드 프루닝을 L1-정규화 최적화 문제의 암묵적 해로 재구성한다.
  • 모든 이전의 임계값 스케줄링 전략이 L1 정규화 항을 조정하는 다양한 방식에 해당함을 규명한다.
  • L1 정규화 계수를 시간에 관계없이 일정하게 유지하는 최적의 임계값 스케줄러를 유도한다. 이는 안정된 목적 함수를 보장한다.
  • 훈련 중 희박화, 조기 프루닝, 초기화 시 프루닝을 포함한 프루닝 방법의 가족을 동일한 프레임워크 내에서 기술한다.
  • ResNet-50, MobileNet-V1, 그리고 SEW ResNet-18에서 프레임워크를 검증하여 모든 설정에서 일관된 성능 향상을 입증한다.
  • 코드를 공개한 LATS(Learned Adaptive Thresholding Scheduler)로 구현한다.

실험 결과

연구 질문

  • RQ1소프트 스레시홀드 프루닝은 기존 최적화 알고리즘인 ISTA와 어떻게 공식적으로 연결될 수 있는가?
  • RQ2이전 소프트 스레시홀드 프루닝 방법에서 사용된 다양한 임계값 스케줄링 전략의 이론적 근거는 무엇인가?
  • RQ3기존 소프트 스레시홀드 프루닝 기법을 설명하고 향상시킬 수 있는 통합 프레임워크를 도출할 수 있는가?
  • RQ4훈련 전반에 걸쳐 일정한 L1 정규화 계수를 유지하는 최적의 임계값 스케줄러를 유도할 수 있는가?
  • RQ5제안된 프레임워크는 다양한 네트워크 아키텍처와 희박성 수준, 스파이킹 신경망까지 일반화 가능한가?

주요 결과

  • 제안된 프레임워크는 소프트 스레시홀드 프루닝이 L1-정규화 최적화 문제에 대한 암묵적 ISTA 솔버임을 입증하며, 기존 방법에 대한 이론적 기반을 제공한다.
  • 모든 이전의 임계값 스케줄링 전략이 L1 정규화 항을 조정하는 다양한 형태임을 규명하며, 하나의 최적화 관점에서 이들의 행동을 통합한다.
  • 유도된 최적의 임계값 스케줄러는 일정한 L1 정규화 계수를 유지함으로써 시간 불변 목적 함수를 확보하고 일반화 성능을 향상시킨다.
  • ImageNet에서 ResNet-50에 대해 99.04%의 희박성에서 88.88%의 Top-1 정확도를 달성하여 최고의 성능을 기록한다.
  • SEW ResNet-18에 대해선 71.18%의 희박성에서 60.11%의 Top-1 정확도, 92.57%의 희박성에서 53.74%의 정확도를 기록하며 기존 방법을 초월한다.
  • 프레임워크를 통해 재학습이나 하이퍼파ram터 검색 없이도 훈련 중 희박화, 조기 프루닝, 초기화 시 프루닝을 포함한 프루닝 전략의 가족을 구현할 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.