Skip to main content
QUICK REVIEW

[논문 리뷰] Stabilizing Transformer Training by Preventing Attention Entropy Collapse

Shuangfei Zhai, Tatiana Likhomanenko|arXiv (Cornell University)|2023. 03. 11.
Advanced Neural Network Applications인용 수 6
한 줄 요약

이 논문은 트랜스포머에서 학습 불안정성의 근본 원인으로 주목받는 주의력 엔트로피 붕괴(자기주의 주의 점수의 병렬적으로 낮은 엔트로피)를 규명한다. 학습 가능한 스칼라를 가진 스펙트럴 정규화 기반의 재구성 방법인 $\sigma$ Reparam을 제안하며, 이는 시각, 음성, 언어 작업 전반에서 학습을 안정화시켜, 웜업, 가중치 감쇠, 적응형 옵timizer 없이도 견고한 학습을 가능하게 한다.

ABSTRACT

Training stability is of great importance to Transformers. In this work, we investigate the training dynamics of Transformers by examining the evolution of the attention layers. In particular, we track the attention entropy for each attention head during the course of training, which is a proxy for model sharpness. We identify a common pattern across different architectures and tasks, where low attention entropy is accompanied by high training instability, which can take the form of oscillating loss or divergence. We denote the pathologically low attention entropy, corresponding to highly concentrated attention scores, as $ extit{entropy collapse}$. As a remedy, we propose $σ$Reparam, a simple and efficient solution where we reparametrize all linear layers with spectral normalization and an additional learned scalar. We demonstrate that $σ$Reparam successfully prevents entropy collapse in the attention layers, promoting more stable training. Additionally, we prove a tight lower bound of the attention entropy, which decreases exponentially fast with the spectral norm of the attention logits, providing additional motivation for our approach. We conduct experiments with $σ$Reparam on image classification, image self-supervised learning, machine translation, speech recognition, and language modeling tasks. We show that $σ$Reparam provides stability and robustness with respect to the choice of hyperparameters, going so far as enabling training (a) a Vision Transformer {to competitive performance} without warmup, weight decay, layer normalization or adaptive optimizers; (b) deep architectures in machine translation and (c) speech recognition to competitive performance without warmup and adaptive optimizers. Code is available at \url{https://github.com/apple/ml-sigma-reparam}.

연구 동기 및 목표

  • 주목력 엔트로피 동역학을 분석함으로써 트랜스포머에서의 학습 불안정성의 근본 원인을 규명한다.
  • 낮은 주목력 엔트로피가 손실의 진동 또는 발산과 관련이 있는 공통 패턴을 식별한다.
  • 다양한 아키텍처와 작업 전반에서 주목력 엔트로피 붕괴를 방지할 수 있는 일반적이고 효율적인 해결책을 개발한다.
  • $\sigma$ Reparam이 최소한의 하이퍼파라미터 튜닝으로 안정적인 학습을 가능하게 함을 보여준다.

제안 방법

  • 모델의 날카움을 측정하는 프록시로 주목력 엔트로피를 주목력 헤드, 레이어, 학습 스텝 전반에 걸쳐 추적한다.
  • $\sigma$ Reparam 도입: 모든 선형 레이어를 스펙트럴 정규화와 학습 가능한 스칼라 $\sigma$를 사용해 재구성한다.
  • 자기주의 주목력 메커니즘의 쿼리, 키, 밸류 프로젝션에 $\sigma$ Reparam을 적용한다.
  • 주목력 엔트로피에 대해 스펙트럴 노름이 증가할수록 지수적으로 감소하는 날카운 빈도 하한을 증명한다.
  • 온도 스케일링 등의 통제된 간섭을 통해 엔트로피 붕괴와 학습 불안정성 간의 인과 관계를 확립한다.
  • $\sigma$ Reparam을 이미지 분류, 자기지도 학습, 기계 번역, 음성 인식, 언어 모델링 등 다양한 작업에서 평가한다.
Figure 1: Transformers are sensitive to hyperparameters. Increasing the learning rate easily causes attention entropy collapse and training divergence. Left: baseline Vision Transformer with default hyperparameters from Touvron et al. ( 2021 ) ; right: $2\times$ learning rate ( $5\times 10^{-4}\maps
Figure 1: Transformers are sensitive to hyperparameters. Increasing the learning rate easily causes attention entropy collapse and training divergence. Left: baseline Vision Transformer with default hyperparameters from Touvron et al. ( 2021 ) ; right: $2\times$ learning rate ( $5\times 10^{-4}\maps

실험 결과

연구 질문

  • RQ1트랜스포머에서 낮은 주목력 엔트로피와 학습 불안정성 간에 일관된 상관관계가 존재하는가?
  • RQ2주목력 엔트로피 붕괴가 손실의 발산 또는 진동과 인과적으로 연결될 수 있는가?
  • RQ3$\sigma$ Reparam이 다양한 작업 전반에서 엔트로피 붕괴를 효과적으로 방지하고 학습을 안정화시키는가?
  • RQ4$\sigma$ Reparam이 가중치 감쇠, 웜업, 적응형 옵티마이저와 같은 일반적인 정규화 기법의 필요성을 제거할 수 있는가?
  • RQ5주목력 로짓의 스펙트럴 노름과 주목력 엔트로피 간의 이론적 관계는 무엇인가?

주요 결과

  • 주목력 엔트로피 붕괴(주목력 점수에서 근접한 0의 엔트로피)는 ViT, NLP, ASR 모델에서 학습 발산 또는 진동과 동시에 발생한다.
  • $\sigma$ Reparam은 엔트로피 붕괴를 성공적으로 방지하여, 웜업, 가중치 감쇠, 레이어 정규화, 또는 적응형 옵티마이저 없이도 비전 트랜스포머의 안정적 학습을 가능하게 한다.
  • 기계 번역 작업에서는 $\sigma$ Reparam 덕분에 이전에 수렴하지 못하던 깊은 트랜스포머 아키텍처의 학습이 가능해졌다.
  • 음성 인식 작업에서는 $\sigma$ Reparam이 웜업이나 적응형 옵티마이저 없이도 기준 성능에 도달하며 경쟁적인 성능을 달성했다.
  • 이론적 분석을 통해 주목력 엔트로피는 스펙트럴 노름이 증가함에 따라 지수적으로 감소하는 날카운 하한을 가짐을 밝혔다.
  • 실험 결과 $\sigma$ Reparam은 하이퍼파라미터 선택에 대해 강건하며, 시각, 음성, 언어 작업 전반에서 단순화된 학습 레시피를 가능하게 했다.
Figure 2: Training can become unstable due to rapid change in attention logit magnitude. We train a Vision Transformer, sharply reducing its temperature in the attention logits by $10\times$ at different intervention epochs. (Blue) Intervention during warmup – at epoch 10 – induces a sharp drop in t
Figure 2: Training can become unstable due to rapid change in attention logit magnitude. We train a Vision Transformer, sharply reducing its temperature in the attention logits by $10\times$ at different intervention epochs. (Blue) Intervention during warmup – at epoch 10 – induces a sharp drop in t

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.