Skip to main content
QUICK REVIEW

[論文レビュー] How to Center Binary Deep Boltzmann Machines

Jan Melchior, Asja Fischer|arXiv (Cornell University)|Nov 6, 2013
Generative Adversarial Networks and Image Synthesis参考文献 17被引用数 3
ひとこと要約

本論文は、可視ユニットおよび隠れユニットの平均値を差し引くことで、二値のディープボルツマンマシン(DBMs)および制限付きボルツマンマシン(RBMs)にセンター化技術を導入する。この手法により、学習の安定性が向上し、対数尤度性能が向上し、自然勾配との更新方向の整合性が高まる。本手法により、グリーディな事前学習の必要性が排除され、指数的移動平均を用いたオフセット推定を組み合わせることで、生成モデルにおいて従来のRBMs/DBMsや強化勾配法を上回る性能を発揮する。

ABSTRACT

This work analyzes centered binary Restricted Boltzmann Machines (RBMs) and binary Deep Boltzmann Machines (DBMs), where centering is done by subtracting offset values from visible and hidden variables. We show analytically that (i) centering results in a different but equivalent parameterization for artificial neural networks in general, (ii) the expected performance of centered binary RBMs/DBMs is invariant under simultaneous flip of data and offsets, for any offset value in the range of zero to one, (iii) centering can be reformulated as a different update rule for normal binary RBMs/DBMs, and (iv) using the enhanced gradient is equivalent to setting the offset values to the average over model and data mean. Furthermore, numerical simulations suggest that (i) optimal generative performance is achieved by subtracting mean values from visible as well as hidden variables, (ii) centered RBMs/DBMs reach significantly higher log-likelihood values than normal binary RBMs/DBMs, (iii) centering variants whose offsets depend on the model mean, like the enhanced gradient, suffer from severe divergence problems, (iv) learning is stabilized if an exponentially moving average over the batch means is used for the offset values instead of the current batch mean, which also prevents the enhanced gradient from diverging, (v) centered RBMs/DBMs reach higher LL values than normal RBMs/DBMs while having a smaller norm of the weight matrix, (vi) centering leads to an update direction that is closer to the natural gradient and that the natural gradient is extremly efficient for training RBMs, (vii) centering dispense the need for greedy layer-wise pre-training of DBMs, (viii) furthermore we show that pre-training often even worsen the results independently whether centering is used or not, and (ix) centering is also beneficial for auto encoders.

研究の動機と目的

  • データのビット反転変換に対して不変性が欠如するRBMs/DBMsの学習における問題を解決すること。
  • 可視ユニットおよび隠れユニットのセンター化が、学習安定性およびモデル性能に与える影響を調査すること。
  • センター化と二値RBMsおよびDBMsにおける自然勾配との関係を分析すること。
  • センター化が、ディープボルツマンマシンにおけるグリーディな階層的事前学習の必要性を排除できるかどうかを評価すること。
  • センター化が、オートエンコーダーおよび生成モデルの性能に与える影響を評価すること。

提案手法

  • センター化は、RBMsおよびDBMsの可視変数および隠れ変数からオフセット値(可視および隠れユニットの平均)を差し引くことで実装される。
  • 本手法は、モデルの同等性を保ちつつ最適化特性を向上させる、等価なパrameterizationに再定式化する。
  • 著者らは、オフセットをデータ平均とモデル平均の平均に設定した場合、センター化と強化勾配法との間で解析的同等性を導出する。
  • オフセット値の推定に、バッチ平均の指数的移動平均が用いられ、学習の発散を防止する。
  • 理論的分析により、センター化が、多様体ベース最適化において最適な自然勾配に近い更新方向をもたらすことが示される。
  • 数値シミュレーションにより、さまざまな学習プロトコル下で、センター化モデルと標準RBMs/DBMs、および強化勾配バージョンとを比較する。

実験結果

リサーチクエスチョン

  • RQ1二値RBMsおよびDBMsにおける可視ユニットおよび隠れユニットのセンター化は、学習安定性および対数尤度性能を向上させるか?
  • RQ2センター化RBMs/DBMsにおける更新方向は、標準モデルと比較して自然勾配に近いか?
  • RQ3センター化により、DBMsにおけるグリーディな階層的事前学習の必要性を排除できるか?
  • RQ4収束性および生成性能の観点から、センター化は強化勾配法と比較して優れているか?
  • RQ5センター化におけるオフセット値推定の最適戦略は何か—バッチ平均、移動平均、またはモデル平均か?

主な発見

  • センター化RBMsおよびDBMsは、重み行列のノルムが小さいにもかかわらず、標準RBMsおよびDBMsよりも顕著に高い対数尤度値を達成する。
  • センター化は学習を安定化させ、特にバッチ平均ではなく指数的移動平均を用いたオフセット推定がなされた場合、発散を防止する。
  • 強化勾配法は、オフセットをデータ平均とモデル平均の平均に設定した場合に等価であるが、適切なオフセット推定がなければ深刻な発散を引き起こす。
  • センター化モデルは、初期学習率やデータ表現にあまり依存せず、標準モデルよりも高い対数尤度値に到達する。
  • センター化により、DBMsにおけるグリーディな事前学習の必要性が排除され、事前学習はセンター化の有無に関わらず性能を低下させることが示された。
  • センター化は、オートエンコーダーにおいても性能を向上させ、生成モデルにとどまらない広範な適用可能性を示している。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。