Skip to main content
QUICK REVIEW

[論文レビュー] SERF: Towards better training of deep neural networks using log-Softplus ERror activation Function

Sayan Nag, Mayukh Bhattacharyya|arXiv (Cornell University)|Aug 21, 2021
Machine Learning in Materials Science被引用数 5
ひとこと要約

本稿では、Swishファミリーから導出された非単調で自己正則化型の活性化関数であるSerfを提案する。この関数は、死滅ReLU問題を軽減し、最適化を改善することを目的としている。Serfは、画像分類、物体検出、機械翻訳、マルチモーダル帰納といった多様なディープラーニングタスクにおいて、ReLU、Swish、Mishを上回る性能を示し、特に深層モデルにおいて一貫した向上を達成している。

ABSTRACT

Activation functions play a pivotal role in determining the training dynamics and neural network performance. The widely adopted activation function ReLU despite being simple and effective has few disadvantages including the Dying ReLU problem. In order to tackle such problems, we propose a novel activation function called Serf which is self-regularized and nonmonotonic in nature. Like Mish, Serf also belongs to the Swish family of functions. Based on several experiments on computer vision (image classification and object detection) and natural language processing (machine translation, sentiment classification and multimodal entailment) tasks with different state-of-the-art architectures, it is observed that Serf vastly outperforms ReLU (baseline) and other activation functions including both Swish and Mish, with a markedly bigger margin on deeper architectures. Ablation studies further demonstrate that Serf based architectures perform better than those of Swish and Mish in varying scenarios, validating the effectiveness and compatibility of Serf with varying depth, complexity, optimizers, learning rates, batch sizes, initializers and dropout rates. Finally, we investigate the mathematical relation between Swish and Serf, thereby showing the impact of preconditioner function ingrained in the first derivative of Serf which provides a regularization effect making gradients smoother and optimization faster.

研究の動機と目的

  • 負の活性化領域におけるゼロ勾配飽和によって引き起こされる死滅ReLU問題を解消すること。
  • 滑らかで前処理付きの活性化関数を通じて、深層ネットワークにおける最適化の安定性と勾配の流れを向上させること。
  • SwishおよびMishの自己ゲーティング機構にインspiredされた自己正則化型活性化関数の開発。
  • 特に深層モデルにおいて、多様なアーキテクチャおよびタスクにおいて優れた性能を示すことを実証すること。
  • 前処理関数がSerfの導関数において果たす数学的役割が正則化の源であることを解明すること。

提案手法

  • Serfを $ f(x) = x \cdot \operatorname{erf}(\ln(1 + e^x)) $ として提案し、自己ゲーティングと誤差関数に基づく非線形性を統合する。
  • Serfの1階微分を活用し、勾配をなめらかにし最適化を促進する前処理関数を埋め込む。
  • ResNet、EfficientNet、YOLOv4、Transformerエンコーダー、BERTベースのアーキテクチャを含む最先端モデルにSerfを統合する。
  • MNISTおよびCIFAR-10におけるアブレーションスタディを通じて、学習率、バッチサイズ、ドロップアウト、重み初期化の変更に伴う性能を評価する。
  • 標準指標(正解率、BLEU、F1スコア)を用いて、複数のベンチマークでReLU、GELU、Swish、MishとSerfを比較する。
  • SwishとSerfの数学的関係を分析し、Serfの導関数における前処理関数の正則化効果を強調する。
Figure 1: Activation functions (Left), first derivatives (Middle) and second derivatives (Right) for Swish, Mish and Serf.
Figure 1: Activation functions (Left), first derivatives (Middle) and second derivatives (Right) for Swish, Mish and Serf.

実験結果

リサーチクエスチョン

  • RQ1Serfは、深層ネットワークにおいて、ReLU、Swish、Mishよりも死滅ReLU問題をより効果的に軽減するか?
  • RQ2Serfの導関数に内蔵された前処理関数は、滑らかな勾配と高速な最適化にどのように寄与するか?
  • RQ3Serfは、多様なアーキテクチャおよびタスクにおいて、特に深層モデルにおいて、どの程度性能を向上させるか?
  • RQ4学習率、バッチサイズ、ドロップアウト率などの変動するトレーニングハイパーパramータ下で、Serfはどの程度の性能を示すか?
  • RQ5Serfは、画像分類、物体検出、機械翻訳、マルチモーダル帰納といった視覚およびNLPタスクに一般化できるか?

主な発見

  • マルチモーダル翻訳タスク(Multi30kドイツ語-英語)において、Serfは36.06のBLEUスコアを達成し、ReLU(35.55)、GELU(35.62)、Mish(35.36)を上回った。
  • IMDb映画レビュー感情分析データセットにおいて、4層のTransformerを用いた場合、Serfは89.03%のトップ1正解率を達成し、ReLU(88.82%)とMish(88.99%)を上回った。
  • PolEmo 2.0感情分析データセットでは、SerfはF1スコア0.8342を達成し、Mishの0.8346にわずかに劣るが、同等の精度と再現率を示した。
  • マルチモーダル帰納タスクにおいて、Serfは5回の実行平均で85.42%の正答率を達成し、GELUの85.28%をわずかに上回った。
  • CIFAR-10およびMNISTにおけるアブレーションスタディにより、Serfが学習率、バッチサイズ、ドロップアウト率、重み初期化手法の変更に対しても頑健であることが確認された。
  • すべての評価タスクおよびアーキテクチャにおいて、SerfはSwishおよびMishを一貫して上回り、特に深層モデルでは顕著な性能差が観察された。
Figure 2: Output landscapes of a randomly initialized 6-layered neural network with ReLU (Left) and Serf (Right) activations.
Figure 2: Output landscapes of a randomly initialized 6-layered neural network with ReLU (Left) and Serf (Right) activations.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。