Skip to main content
QUICK REVIEW

[論文レビュー] A U-turn on Double Descent: Rethinking Parameter Counting in Statistical Learning

Alicia Curth, Alan Jeffares|arXiv (Cornell University)|Oct 29, 2023
Neural Networks and Applications被引用数 7
ひとこと要約

この論文は、古典的機械学習モデルにおける二重降下の解釈に挑戦し、テスト誤差の見かけ上の第二の降下が複数の複雑さ軸を混同することに起因する誤りであることを示している。非パラメトリックスムージングに基づく有効パラメータ数を導入することで、複雑さが適切に測定されると、テスト誤差曲線は従来のU字型の形に戻ることを実証し、二重降下と古典的統計理論の間の矛盾を解消している。

ABSTRACT

Conventional statistical wisdom established a well-understood relationship between model complexity and prediction error, typically presented as a U-shaped curve reflecting a transition between under- and overfitting regimes. However, motivated by the success of overparametrized neural networks, recent influential work has suggested this theory to be generally incomplete, introducing an additional regime that exhibits a second descent in test error as the parameter count p grows past sample size n - a phenomenon dubbed double descent. While most attention has naturally been given to the deep-learning setting, double descent was shown to emerge more generally across non-neural models: known cases include linear regression, trees, and boosting. In this work, we take a closer look at evidence surrounding these more classical statistical machine learning methods and challenge the claim that observed cases of double descent truly extend the limits of a traditional U-shaped complexity-generalization curve therein. We show that once careful consideration is given to what is being plotted on the x-axes of their double descent plots, it becomes apparent that there are implicitly multiple complexity axes along which the parameter count grows. We demonstrate that the second descent appears exactly (and only) when and where the transition between these underlying axes occurs, and that its location is thus not inherently tied to the interpolation threshold p=n. We then gain further insight by adopting a classical nonparametric statistics perspective. We interpret the investigated methods as smoothers and propose a generalized measure for the effective number of parameters they use on unseen examples, using which we find that their apparent double descent curves indeed fold back into more traditional convex shapes - providing a resolution to tensions between double descent and statistical intuition.

研究の動機と目的

  • 決定木、ブースティング、線形回帰などの非ニューラルモデルにおける二重降下が、古典的U字型一般化曲線を破壊するという主張に反論すること。
  • 二重降下のプロットが、複数の複雑さ軸を暗黙的に組み合わせており、単一の合成軸に沿ってプロットすると誤解を招くことの特定。
  • スムージング器における一般化された有効パラメータ数を提案し、未知データにおけるモデルの複雑さを測定すること。
  • 二重降下が $p = n$ の補間閾値に固有のものではなく、軸の遷移に起因する誤りであることを示すこと。
  • 一般化誤差を有効自由度の観点から再表現することで、二重降下と古典的統計的直観を調和させること。

提案手法

  • 決定木、ブースティング、線形モデルをスムージング器として解釈する非パラメトリックな視点を導入する。
  • スムージング行列のトレースを用いて、テスト入力が予測に与える影響に基づき、有効パラメータ数 $p^{ ext{test}}_{ ilde{ extbf{s}}}$ を定義する。
  • 一般化プロットのx軸として、生のモデルパラメータやハイパーパrameterではなく、この有効パラメータ数を用いる。
  • 木の深さや推定器数などの別々のハイパーパrameter軸に沿って分解することで、決定木およびブースティングモデルにおける二重降下を分析する。
  • テスト誤差を $p^{ ext{test}}_{ ilde{ extbf{s}}}$ に沿って再構築することで、第二の降下が消失し、曲線が凸型になることを示す。
  • 線形回帰に最小ノルム解を適用し、非監視次元削減と有効自由度を関連付ける。

実験結果

リサーチクエスチョン

  • RQ1深層学習以外のモデルにおける観察された二重降下は、本当に古典的U字型一般化曲線からの逸脱であるのか?
  • RQ2パラメータ数が標本サイズを超えた際に、二重降下プロットにおける第二の降下は何かの原因によって生じるのか?
  • RQ3二重降下が複数の複雑さ軸を組み合わせた結果であるに過ぎず、モデル行動の根本的転換ではないという説明は可能か?
  • RQ4有効パラメータ数による複雑さの測定によって、これらのモデルにおける古典的U字型曲線が回復するのか?
  • RQ5有効パラメータ数は、異なるモデルクラスにおける補間閾値 $p = n$ とどのように関係するのか?

主な発見

  • 決定木およびブースティングモデルにおける二重降下は、木の深さや推定器数といった複数のハイパーパrameterを段階的に増加させることに起因し、単一の複雑さ軸によるものではない。
  • 別々の軸に沿ってプロットすると、木の深さおよび推定器数の両方が古典的U字型曲線を示し、古典理論からの根本的逸脱は示さない。
  • 二重降下プロットにおける第二の降下は、これらの基本的複雑さ軸の遷移点に正確に一致しており、$p = n$ に固有のものではない。
  • 最小ノルム解を用いた線形回帰では、有効パラメータ数 $p^{ ext{test}}_{ ilde{ extbf{s}}}$ は有界であり、$p = n$ の補間閾値を超えて増加しない。
  • テスト誤差を $p^{ ext{test}}_{ ilde{ extbf{s}}}$ に沿ってプロットすると、決定木、ブースティング、RFF回帰を含むすべてのモデルで二重降下曲線が凸型のU字型曲線に収束する。
  • 訓練誤差がゼロである補間モデルの最良性能を示すものは、常に $p^{ ext{test}}_{ ilde{ extbf{s}}}$ が最小のものであり、有効パラメータ数の予測力が裏付けられる。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。