Skip to main content
QUICK REVIEW

[논문 리뷰] SERF: Towards better training of deep neural networks using log-Softplus ERror activation Function

Sayan Nag, Mayukh Bhattacharyya|arXiv (Cornell University)|2021. 08. 21.
Machine Learning in Materials Science인용 수 5
한 줄 요약

이 논문은 스위시 가족에서 유도된 비단조화적이고 자기정규화된 활성화 함수인 Serf를 제안한다. 이는 사라지는 ReLU 문제를 완화하고 최적화를 향상시키기 위해 고안되었다. Serf는 이미지 분류, 객체 검출, 기계 번역, 다중모odal 함의 등 다양한 딥러닝 작업에서 ReLU, 스위시, 미시보다 뛰어난 성능을 보이며, 깊이 있는 아키텍처에서 일관된 성능 향상을 보였다.

ABSTRACT

Activation functions play a pivotal role in determining the training dynamics and neural network performance. The widely adopted activation function ReLU despite being simple and effective has few disadvantages including the Dying ReLU problem. In order to tackle such problems, we propose a novel activation function called Serf which is self-regularized and nonmonotonic in nature. Like Mish, Serf also belongs to the Swish family of functions. Based on several experiments on computer vision (image classification and object detection) and natural language processing (machine translation, sentiment classification and multimodal entailment) tasks with different state-of-the-art architectures, it is observed that Serf vastly outperforms ReLU (baseline) and other activation functions including both Swish and Mish, with a markedly bigger margin on deeper architectures. Ablation studies further demonstrate that Serf based architectures perform better than those of Swish and Mish in varying scenarios, validating the effectiveness and compatibility of Serf with varying depth, complexity, optimizers, learning rates, batch sizes, initializers and dropout rates. Finally, we investigate the mathematical relation between Swish and Serf, thereby showing the impact of preconditioner function ingrained in the first derivative of Serf which provides a regularization effect making gradients smoother and optimization faster.

연구 동기 및 목표

  • 음성 활성화 영역에서 0 기울기 포화 상태로 인해 발생하는 사라지는 ReLU 문제를 해결하기 위해.
  • 부드럽고 조정된 활성화 함수를 통해 깊은 신경망에서 최적화 안정성과 기울기 흐름을 향상시키기 위해.
  • 스위시와 미시의 자기게이팅 메커니즘에서 영감을 얻은 자기정규화 활성화 함수를 개발하기 위해.
  • 특히 깊은 모델에서 성능 향상이 뚜렷한 다양한 아키텍처와 작업에서 뛰어난 성능을 입증하기 위해.
  • 조정자(preconditioner)가 Serf의 도함수에서 정규화 효과를 어떻게 유도하는지 수학적으로 분석하기 위해.

제안 방법

  • Serf를 $ f(x) = x \cdot \operatorname{erf}(\ln(1 + e^x)) $ 로 제안하여 자기게이팅과 오차함수 기반 비선형성의 조합을 구현한다.
  • Serf의 첫 번째 도함수를 활용해 기울기의 매끄러움과 최적화 향상을 향상시키는 조정자 함수를 통합한다.
  • ResNet, EfficientNet, YOLOv4, 트랜스포머 인코더, BERT 기반 아키텍처를 포함한 최첨단 모델들에 Serf를 통합한다.
  • 학습률, 배치 크기, 드롭아웃, 가중치 초기화 방법 등 다양한 초모델 설정에서 성능을 평가하기 위해 MNIST와 CIFAR-10에서 분석 실험을 수행한다.
  • 표준 평가 지표(정확도, BLEU, F1 점수)를 사용해 여러 벤치마크에서 ReLU, GELU, 스위시, 미시와 Serf를 비교한다.
  • 스위시와 Serf 간의 수학적 관계를 분석하여 Serf 도함수 내 조정자의 정규화 효과를 부각시킨다.
Figure 1: Activation functions (Left), first derivatives (Middle) and second derivatives (Right) for Swish, Mish and Serf.
Figure 1: Activation functions (Left), first derivatives (Middle) and second derivatives (Right) for Swish, Mish and Serf.

실험 결과

연구 질문

  • RQ1Serf는 깊은 신경망에서 ReLU, 스위시, 미시보다 사라지는 ReLU 문제를 더 효과적으로 완화하는가?
  • RQ2Serf 도함수 내 조정자가 기울기의 매끄러움과 더 빠른 최적화에 어떻게 기여하는가?
  • RQ3Serf는 다양한 아키텍처와 작업에서 성능 향상을 얼마나 높게 보이는가, 특히 깊은 모델에서의 성능 향상은 어떠한가?
  • RQ4Serf는 학습 초모델 설정(학습률, 배치 크기, 드롭아웃 비율 등)이 다양할 경우 어떻게 성능을 보이는가?
  • RQ5Serf는 시각 및 자연어처리 작업 전반에 걸쳐 일반화 가능한가, 즉 이미지 분류, 객체 검출, 기계 번역, 다중모달 함의 작업 등에서 성능이 우수한가?

주요 결과

  • 멀티30K 독일-영어 번역 작업에서 Serf는 BLEU 점수 36.06을 기록하여 ReLU(35.55), GELU(35.62), 미시(35.36)를 모두 초월했다.
  • IMDb 영화 리뷰 감성 분석 데이터셋에서 4층 트랜스포머를 사용한 Serf는 89.03%의 상위 1위 정확도를 기록하여 ReLU(88.82%)와 미시(88.99%)를 뛰어넘었다.
  • PolEmo 2.0 감성 분석 데이터셋에서 Serf는 F1 점수 0.8342를 기록했으며, 미시의 0.8346과 비슷한 정밀도와 재현율을 보였다.
  • 다중모달 함의 작업에서 Serf는 5회 반복 평균 정확도 85.42%를 기록하여 GELU의 85.28%를 약간 상회했다.
  • CIFAR-10과 MNIST에서의 분석 실험을 통해 Serf가 다양한 학습률, 배치 크기, 드롭아웃 비율, 가중치 초기화 방법 설정에서 뛰어난 강건성을 보였다.
  • 모든 평가된 작업과 아키텍처에서 Serf는 스위시와 미시를 지속적으로 능가했으며, 특히 깊은 모델에서 더 큰 성능 격차를 보였다.
Figure 2: Output landscapes of a randomly initialized 6-layered neural network with ReLU (Left) and Serf (Right) activations.
Figure 2: Output landscapes of a randomly initialized 6-layered neural network with ReLU (Left) and Serf (Right) activations.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.