Skip to main content
QUICK REVIEW

[論文レビュー] Theoretical foundation for CMA-ES from information geometric perspective

Youhei Akimoto, Yuichi Nagata|arXiv (Cornell University)|Jun 4, 2012
Face and Expression Recognition参考文献 21被引用数 4
ひとこと要約

この論文は情報幾何を用いて、共分散行列適応進化戦略(CMA-ES)の理論的基盤を確立し、そのパrameter更新が期待適応度を最大化するための自然勾配学習に等価であることを示している。主な結果は、CMA-ESがフィッシャー情報行列の逆行列を明示的に計算することなく、自然勾配上昇を暗黙的に行っていることであり、これによりそのデフォルトの学習率が正当化され、一般化期待最大化フレームワークと結びつく。

ABSTRACT

This paper explores the theoretical basis of the covariance matrix adaptation evolution strategy (CMA-ES) from the information geometry viewpoint. To establish a theoretical foundation for the CMA-ES, we focus on a geometric structure of a Riemannian manifold of probability distributions equipped with the Fisher metric. We define a function on the manifold which is the expectation of fitness over the sampling distribution, and regard the goal of update of the parameters of sampling distribution in the CMA-ES as maximization of the expected fitness. We investigate the steepest ascent learning for the expected fitness maximization, where the steepest ascent direction is given by the natural gradient, which is the product of the inverse of the Fisher information matrix and the conventional gradient of the function. Our first result is that we can obtain under some types of parameterization of multivariate normal distribution the natural gradient of the expected fitness without the need for inversion of the Fisher information matrix. We find that the update of the distribution parameters in the CMA-ES is the same as natural gradient learning for expected fitness maximization. Our second result is that we derive the range of learning rates such that a step in the direction of the exact natural gradient improves the parameters in the expected fitness. We see from the close relation between the CMA-ES and natural gradient learning that the default setting of learning rates in the CMA-ES seems suitable in terms of monotone improvement in expected fitness. Then, we discuss the relation to the expectation-maximization framework and provide an information geometric interpretation of the CMA-ES.

研究の動機と目的

  • 情報幾何を用いてCMA-ESアルゴリズムの理論的裏付けを提供すること。
  • 多変量正規分布のリーマン多様体上でのCMA-ESと自然勾配学習の関係を調査すること。
  • CMA-ESのデフォルトの学習率が、期待適応度の単調な改善を保証するためになぜ有効であるかを説明すること。
  • 情報幾何的視点からCMA-ESを一般化期待最大化(GEM)アルゴリズムとして解釈すること。
  • CMA-ESと自然勾配法との間の接続を、特にパrameter更新ダイナミクスの観点から確立すること。

提案手法

  • 著者たちは、フィッシャー計量を備えたリーマン多様体上に、多変量正規分布としてのサンプリング分布をモデル化する。
  • 期待適応度をこの多様体上の関数として定義し、自然勾配による勾配上昇を分析する。自然勾配は、フィッシャー情報行列の逆行列と通常の勾配の積として定義される。
  • 多変量正規分布の特定のパrameter化のもとで、フィッシャー情報行列の逆行列を明示的に計算せずに、期待適応度の自然勾配が計算可能であることを示す。
  • CMA-ESのパrameter更新ルールが、サンプルベースの自然勾配推定を用いた自然勾配学習と数学的に同等であることを示す。
  • 正確な自然勾配方向に従う場合に、期待適応度の単調な改善を保証する学習率の範囲を導出する。
  • CMA-ESを期待最大化(EM)フレームワークと比較し、KL発散と適応度改善の両方の項を考慮する一般化EM(GEM)アルゴリズムとして解釈する。

実験結果

リサーチクエスチョン

  • RQ1CMA-ESのパrameter更新ルールは、多変量正規分布の多様体上での自然勾配学習とどのように関係しているか?
  • RQ2特定のパrameter化のもとで、フィッシャー情報行列の逆行列を明示的に計算せずに、期待適応度の自然勾配を計算できるか?
  • RQ3正確な自然勾配方向に従う場合に、期待適忺度の単調な改善を保証する学習率の範囲は何か?
  • RQ4情報幾何的構造と収束特性の観点から、CMA-ESはEMおよびGEMフレームワークとどのように比較できるか?
  • RQ5適応度重み付きサンプリングで定義されるターゲット分布との関係において、CMA-ESの情報幾何的解釈は何か?

主な発見

  • CMA-ESのパrameter更新ルールは、サンプリング分布下での期待適応度を最大化するための自然勾配学習と数学的に同等である。
  • 多変量正規分布の特定のパrameter化のもとで、フィッシャー情報行列の逆行列を計算せずに、期待適応度の自然勾配をサンプルから直接推定できる。
  • CMA-ESで使用されるデフォルトの学習率は、正確な自然勾配方向に従う場合に期待適応度の単調な改善を保証する範囲にあることから正当化される。
  • CMA-ESは、標準のEMが片方の項のみを最大化するのに対し、ターゲット分布へのKL発散と期待適応度の両方を改善する一般化EM(GEM)アルゴリズムとして解釈できる。
  • 情報幾何的視点により、CMA-ESが適応度改善の2つの項をバランスさせていることが明らかになった:サンプリング分布の変化と、適応度重み付きターゲット分布の変化であり、これは標準のEMよりもより強固な更新を提供する。
  • 自然勾配の定式化により、CMA-ESに原理的基盤が与えられ、今後の研究において収束理論の向上やより効率的なパrameter更新のための可能性が示唆される。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。