Skip to main content
QUICK REVIEW

[論文レビュー] General Information Bottleneck Objectives and their Applications to Machine Learning

Sayandev Mukherjee|arXiv (Cornell University)|Dec 12, 2019
Neural Networks and Applications参考文献 12被引用数 4
ひとこと要約

本稿は、情報ボトルネック原理(IBP)と予測情報ボトルネック原理(PIBP)を一般化する統一的な一般情報ボトルネック目的関数(IBO)の族を導入し、IBP/PIBP(モデルパラメータとテストデータの間の相互情報量を最大化することを提唱)と最近の結果(この相互情報量を最小化することを示唆)の間にある矛盾を解消する。モデルパラメータが経験的損失を最小化するために選ばれるという事実を考慮に入れた新たなIBOを導出し、最適化のための変分上界を提供することで、理論的推奨事項の矛盾を解消する。

ABSTRACT

We view the Information Bottleneck Principle (IBP: Tishby et al., 1999; Schwartz-Ziv and Tishby, 2017) and Predictive Information Bottleneck Principle (PIBP: Still et al., 2007; Alemi, 2019) as special cases of a family of general information bottleneck objectives (IBOs). Each IBO corresponds to a particular constrained optimization problem where the constraints apply to: (a) the mutual information between the training data and the learned model parameters or extracted representation of the data, and (b) the mutual information between the learned model parameters or extracted representation of the data and the test data (if any). The heuristics behind the IBP and PIBP are shown to yield different constraints in the corresponding constrained optimization problem formulations. We show how other heuristics lead to a new IBO, different from both the IBP and PIBP, and use the techniques from (Alemi, 2019) to derive and optimize a variational upper bound on the new IBO. We then apply the theory of general IBOs to resolve the seeming contradiction between, on the one hand, the recommendations of IBP and PIBP to maximize the mutual information between the model parameters and test data, and on the other, recent information-theoretic results (see Xu and Raginsky, 2017) suggesting that this mutual information should be minimized. The key insight is that the heuristics (and thus the constraints in the constrained optimization problems) of IBP and PIBP are not applicable to the scenario analyzed by (Xu and Raginsky, 2017) because the latter makes the additional assumption that the parameters of the trained model have been selected to minimize the empirical loss function. Aided by this insight, we formulate a new IBO that accounts for this property of the parameters of the trained model, and derive and optimize a variational bound on this IBO.

研究の動機と目的

  • 情報ボトルネック原理(IBP)と予測IBP(PIBP)が、モデルパラメータとテストデータの間の相互情報量を最大化することを提唱するのに対し、最近の情報理論的結果(XuとRaginsky, 2017)がその相互情報量を最小化することを示唆するという、表面的な矛盾を解消すること。
  • IBPとPIBPを、モデルパラメータとトレーニング/テストデータの間の相互情報量に関する制約を含む、より広範な制約付き最適化問題の族として形式化すること。
  • IBPとPIBPの背後にあるヒューリスティック仮定を特定し、それらがXuとRaginsky(2017)が分析した状況(モデルパラメータが経験的損失を最小化するために選ばれるという仮定)には適用できないことを明らかにすること。
  • モデルパラメータがトレーニングセット上で経験的損失を最小化するために選ばれるという制約を明示的に組み込んだ新たな情報ボトルネック目的関数(IBO)を提案すること。
  • Alemi(2019)の手法を用いて、新しいIBOの変分上界を導出し、機械学習における実用的応用を可能にする。

提案手法

  • IBPとPIBPを、$ I_1 - \nu I_2 $ の形をとる一般的な情報ボトルネック目的関数(IBO)の族として形式化し、ここで $ I_1 $ と $ I_2 $ はそれぞれモデルパラメータとトレーニング/テストデータの間の相互情報量である。
  • IBPとPIBPが、$ I(t;\bm{x_P}) $ と $ I(t;\bm{x_F}) $ に異なる制約を課す、異なる制約付き最適化問題に対応することを特定。これは、それらの背後にあるヒューリスティックに起因する。
  • IBPとPIBPのヒューリスティックが、XuとRaginsky(2017)の状況(モデルパラメータが経験的損失を最小化するために選ばれる)には適用できないことを示す。
  • 新たなIBOを提案:$ \max_{p(\theta|\bm{x_P})} \left[ I(\theta;\bm{x_P}|\bm{x_F}) - \beta I(\theta;\bm{x_P}) \right] $、制約条件として $ L(\theta,\bm{x_P}) \leq \epsilon $、ここで $ \beta \geq 1 $。
  • Alemi(2019)のフレームワークを用いて、新しいIBOの変分上界を導出し、変分推論を用いた最適化を可能にする。
  • 次式の上界を適用:$ I(\theta;\bm{x_P}|\bm{x_F}) - \beta I(\theta;\bm{x_P}) \leq -\beta H(\bm{x_P}) - \mathbb{E}_{p(\phi)p(\bm{x_P}|\phi)} \log Z_\beta(\bm{x_P}) $、ここで $ Z_\beta $ は補助分布 $ q(\theta) $ と $ q(\bm{x_P}|\theta) $ を用いて定義される。

実験結果

リサーチクエスチョン

  • RQ1なぜIBPとPIBPは、モデルパラメータとテストデータの間の相互情報量を最大化することを提唱するのに対し、最近の一般化境界はその最小化を示唆するのか?
  • RQ2IBPとPIBPのヒューリスティックに内在する仮定は何か?それらがなぜXuとRaginsky(2017)の状況に適用できないのか?
  • RQ3モデルパラメータが経験的損失を最小化するために選ばれるという事実を考慮に入れた、新たな情報ボトルネック目的関数をどのように定式化できるか?
  • RQ4この新たなIBOに対して、変分上界を導出し、最適化可能にすることができるか?
  • RQ5新たなIBOは、一般化理論における相互情報量の最大化と最小化という、矛盾する推奨事項をどのように調和させるのか?

主な発見

  • IBPとPIBPが、モデルパラメータとトレーニング/テストデータの間の相互情報量に関する制約を含む、より広範な制約付き最適化問題の特殊ケースであることが示された。
  • 主な洞察は、IBPとPIBPのヒューリスティックが、XuとRaginsky(2017)の状況(モデルパラメータが経験的損失を最小化するために選ばれる)には適用できないことである。この状況では、IBP や PIBP が仮定するようなパラメータ選択のメカニズムが成り立たない。
  • 経験的損失最小化の制約を明示的に組み込んだ新たなIBOを提案:$ \max_{p(\theta|\bm{x_P})} \left[ I(\theta;\bm{x_P}|\bm{x_F}) - \beta I(\theta;\bm{x_P}) \right] $、$ \beta \geq 1 $。
  • 本稿では、新しいIBOのタイトな変分上界を導出し:$ I(\theta;\bm{x_P}|\bm{x_F}) - \beta I(\theta;\bm{x_P}) \leq -\beta H(\bm{x_P}) - \mathbb{E}_{p(\phi)p(\bm{x_P}|\phi)} \log Z_\beta(\bm{x_P}) $、これにより変分推論による最適化が可能になる。
  • 新しいIBOは、モデルパラメータがトレーニングセット上で経験的損失を最小化するために選ばれるという事実を明示的にモデル化することで、IBP/PIBPと最近の一般化境界の間の矛盾を解消する。
  • 提案されたフレームワークは、既存のIBPとPIBPの目的関数を統一的に扱い、現実のトレーニングダイナミクスを反映する新たな目的関数を体系的かつ原理的につくるための手法を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。