Skip to main content
QUICK REVIEW

[論文レビュー] Revisiting complexity and the bias-variance tradeoff.

Raaz Dwivedi, Chandan Singh|arXiv (Cornell University)|Jun 17, 2020
Sparse and Compressive Sensing Techniques参考文献 60被引用数 9
ひとこと要約

この論文は、Rissanenの最小記述長(MDL)原理に基づく新しい複雑さ測度(MDL-COMP)を提案することで、高次元モデルにおけるバイアス-バリアンストレードオフを再考する。MDL-COMPは高次元において log d のスケーリングを示す(d/n よりも遅い)ため、DNNのような適切にチューニングされた高次元モデルの一般化を原理的かつ明確に説明できる。

ABSTRACT

The recent success of high-dimensional models, such as deep neural networks (DNNs), has led many to question the validity of the bias-variance tradeoff principle in high dimensions. We reexamine it with respect to two key choices: the model class and the complexity measure. We argue that failing to suitably specify either one can falsely suggest that the tradeoff does not hold. This observation motivates us to seek a valid complexity measure, defined with respect to a reasonably good class of models. Building on Rissanen's principle of minimum description length (MDL), we propose a novel MDL-based complexity (MDL-COMP). We focus on the context of linear models, which have been recently used as a stylized tractable approximation to DNNs in high-dimensions. MDL-COMP is defined via an optimality criterion over the encodings induced by a good Ridge estimator class. We derive closed-form expressions for MDL-COMP and show that for a dataset with $n$ observations and $d$ parameters it is \emph{not always} equal to $d/n$, and is a function of the singular values of the design matrix and the signal-to-noise ratio. For random Gaussian design, we find that while MDL-COMP scales linearly with $d$ in low-dimensions ($d n$) the scaling is exponentially smaller, scaling as $\log d$. We hope that such a slow growth of complexity in high-dimensions can help shed light on the good generalization performance of several well-tuned high-dimensional models. Moreover, via an array of simulations and real-data experiments, we show that a data-driven Prac-MDL-COMP can inform hyper-parameter tuning for ridge regression in limited data settings, sometimes improving upon cross-validation.

研究の動機と目的

  • 適切に選ばれたモデルクラスに基づくモデルの複雑さの再定義を通じて、高次元モデルにおけるバイアス-バリアンストレードオフを再表現すること。
  • 不適切な複雑さ測度による誤解から生じる、高次元においてバイアス-バリアンストレードオフが崩壊するという誤った印象を是正すること。
  • 線形モデルに対して、最小記述長(MDL)原理に基づく原理的かつデータ駆動型の複雑さ測度を構築すること。
  • MDL-COMPが常に d/n に等しいわけではないこと、および特異値と信号対雑音比に依存することを示すこと。
  • MDL-COMPが限られたデータ環境下でのハイパーパramータチューニングを支援できることを示し、一部の状況では交差検証を上回ることを示すこと。

提案手法

  • RissanenのMDL原理に従い、良好なリッジ推定器クラスによって誘導される符号化の最適性基準に基づき、MDL-COMPを定義する。
  • 設計行列の特異値と信号対雑音比に依存するMDL-COMPの閉形式表現を導出する。
  • ランダムなガウス設計下でのMDL-COMPの分析により、低次元ではdに線形に依存するが、高次元では log d にスケーリングすることを示す。
  • 実用的なハイパーパramータチューニングのためのデータ駆動型バージョン、Prac-MDL-COMPを導入する。
  • 低データレジームにおける交差検証との比較を通じて、シミュレーションおよび実データ実験により手法を検証する。

実験結果

リサーチクエスチョン

  • RQ1適切に定義された複雑さ測度を用いる場合、高次元モデルにおけるバイアス-バリアンストレードオフは依然として有効であるか?
  • RQ2MDL原理から導出された複雑さ測度は、高次元線形モデルにおけるモデル複雑さのより正確な特徴付けを提供できるか?
  • RQ3ランダムなガウス設計下で、MDL-COMPは次元dとともにどのようにスケーリングするか?また、従来のd/n測度とは異なるか?
  • RQ4Prac-MDL-COMPは、データが限られた環境下で交差検証に比べてハイパーパramータチューニングを改善できるか?
  • RQ5設計行列の特異値と信号対雑音比は、MDL-COMPの決定にどのような役割を果たすか?

主な発見

  • MDL-COMPは常に d/n に等しいわけではない。設計行列の特異値と信号対雑音比に依存する。
  • ランダムなガウス設計下では、MDL-COMPは低次元ではdに線形に依存するが、高次元では log d にスケーリングする。
  • 高次元における遅い log d スケーリングは、適切にチューニングされた高次元モデルの優れた一般化性能の潜在的説明である。
  • データ駆動型バージョンであるPrac-MDL-COMPは、限られたデータ下でのリッジ回帰におけるハイパーパramータチューニングを改善でき、一部の状況では交差検証を上回る。
  • 提案された複雑さ測度は、高次元線形モデリングにおけるヒューリスティック的または漸近的定義に代わる原理的代替手段を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。