Skip to main content
QUICK REVIEW

[논문 리뷰] On the Promise of the Stochastic Generalized Gauss-Newton Method for Training DNNs

Matilde Gargiani, Andrea Zanelli|arXiv (Cornell University)|2020. 06. 03.
Stochastic Gradient Optimization Techniques인용 수 5
한 줄 요약

이 논문은 헤시안 행렬을 직접 계산하지 않고 가우스-뉴턴 헤시안 근사와 공액 그래디언트 반복을 사용하여 깊이 신경망을 훈련시키는 헤시안-프리(second-order) 최적화 기법인 확률적 일반화 가우스-뉴턴(SGN) 방법을 제안한다. SGN은 특히 대용량 미니배치 설정에서 SGD보다 수렴 반복 수와 런타임을 크게 감소시키며, 단계 크기와 같은 하이퍼파rameter 조정에 대해 뛰어난 강건성을 보이며, 일阶 최적화 방법의 유망한 대안이 된다.

ABSTRACT

Following early work on Hessian-free methods for deep learning, we study a stochastic generalized Gauss-Newton method (SGN) for training DNNs. SGN is a second-order optimization method, with efficient iterations, that we demonstrate to often require substantially fewer iterations than standard SGD to converge. As the name suggests, SGN uses a Gauss-Newton approximation for the Hessian matrix, and, in order to compute an approximate search direction, relies on the conjugate gradient method combined with forward and reverse automatic differentiation. Despite the success of SGD and its first-order variants, and despite Hessian-free methods based on the Gauss-Newton Hessian approximation having been already theoretically proposed as practical methods for training DNNs, we believe that SGN has a lot of undiscovered and yet not fully displayed potential in big mini-batch scenarios. For this setting, we demonstrate that SGN does not only substantially improve over SGD in terms of the number of iterations, but also in terms of runtime. This is made possible by an efficient, easy-to-use and flexible implementation of SGN we propose in the Theano deep learning platform, which, unlike Tensorflow and Pytorch, supports forward automatic differentiation. This enables researchers to further study and improve this promising optimization technique and hopefully reconsider stochastic second-order methods as competitive optimization techniques for training DNNs; we also hope that the promise of SGN may lead to forward automatic differentiation being added to Tensorflow or Pytorch. Our results also show that in big mini-batch scenarios SGN is more robust than SGD with respect to its hyperparameters (we never had to tune its step-size for our benchmarks!), which eases the expensive process of hyperparameter tuning that is instead crucial for the performance of first-order methods.

연구 동기 및 목표

  • 깊이 신경망 훈련을 위한 확률적 일반화 가우스-뉴턴(SGN) 방법의 경험적 성능과 실용적 타당성을 조사하기 위해.
  • 표준 SGD와 비교하여 SGN의 수렴 속도, 런타임 효율성, 하이퍼파ram터에 대한 강건성을 평가하기 위해.
  • SGN이 반복 수뿐만 아니라 총 훈련 시간에서도 SGD를 능가할 수 있음을 입증하기 위해, 특히 대용량 미니배치 설정에서.
  • PyTorch 및 TensorFlow와 같은 딥러닝 프레임워크에 전방 자동미분(forward automatic differentiation)을 통합하여 고급 두 번째 단계 최적화 기법의 광범위한 채택을 가능하게 하기 위해.
  • 딥러닝에서 첫 번째 단계 최적화의 실용적이고 확장 가능한 대안으로서, 확률적 두 번째 단계 최적화 방법에 대한 관심을 재진입시키기 위해.

제안 방법

  • SGN은 헤시안 행렬의 가우스-뉴턴 근사를 사용하여 명시적 헤시안 계산 없이도 곡률 정보를 효율적으로 포착한다.
  • 결과로 생기는 선형 시스템을 공액 그래디언트(CG) 방법을 통해 해결하며, 헤시안-벡터 곱을 계산하기 위해 전방 및 역방향 자동미분을 활용한다.
  • 계산 비용을 줄이기 위해 확률적 미니배치를 사용하면서도 수렴 보장을 유지한다.
  • Theano에서의 구현은 전방 자동미분을 지원하여 곡률 연산을 효율적이고 유연하게 계산할 수 있도록 한다.
  • 잔차 크기에 기반해 CG 반복 수를 동적으로 조정하며, 대용량 미니배치 환경에서는 더 높은 반복 수를 권장하는 히우리스틱을 제안한다.
  • 최적화 중 CG 반복 수를 조정하기 위한 적응형 신뢰역(trust-region) 기법을 지원하여 수렴 안정성을 향상시킨다.

실험 결과

연구 질문

  • RQ1확률적 일반화 가우스-뉴턴(SGN) 방법이 깊이 신경망 훈련에서 반복 수와 런타임 측면에서 SGD보다 더 빠른 수렴을 달성할 수 있는가?
  • RQ2계산 효율성과 하이퍼파ram터에 대한 강건성이 핵심이 되는 대용량 미니배치 환경에서 SGN의 성능은 어떠한가?
  • RQ3SGD에 비해 SGN이 하이퍼파ram터 조정, 특히 단계 크기에 대해 얼마나 강건한가?
  • RQ4Theano, PyTorch 또는 TensorFlow와 같은 딥러닝 프레임워크에 전방 자동미분를 통합함으로써 고급 두 번째 단계 최적화 기법의 광범위한 채택이 가능해지는가?
  • RQ5시간 제약이 있는 훈련 환경에서 성능을 극대화하기 위해 SGN의 CG 반복 수와 계산 비용 사이의 최적의 트레이드오프는 무엇인가?

주요 결과

  • SGN은 모든 벤치마크에서 특히 대용량 미니배치 설정에서 SGD보다 유의미하게 적은 반복 수로 수렴한다.
  • 더 비싼 반복를 거치더라도 SGN은 더 빠른 수렴 덕분에 경쟁력 있거나 더 낫게 런타임 성능을 보였다.
  • SGN은 하이퍼파ram터 조정, 특히 단계 크기에 대해 놀라운 강건성을 보였으며, 반면 SGD는 학습률 값에 매우 민감하게 반응했다.
  • SGN은 일반화 성능가 더 뛰어나, 더 적은 에포크 수로도 더 높은 테스트 정확도를 달성하여 향상된 일반화 성질을 나타냈다.
  • Theano에서의 전방 자동미분 사용은 SGN의 효율적이고 민감한 구현을 가능하게 하였으며, 적응형 CG 반복 제어와 같은 고급 기능을 지원했다.
  • 결과는 SGN이 광범위한 하이퍼파라미터 튜닝의 필요성을 줄일 수 있으며, 더 지속 가능하고 에너지 효율적인 딥러닝 훈련에 기여할 수 있음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.