Skip to main content
QUICK REVIEW

[论文解读] Bayes-optimal Learning of Deep Random Networks of Extensive-width

Hugo Cui, Florent Krząkała|arXiv (Cornell University)|Feb 1, 2023
Machine Learning and Algorithms被引用 6
一句话总结

该论文通过一个新颖的高斯等价性性质(GEP)猜想,推导出具有高斯输入的深度广义宽度随机神经网络的贝叶斯最优测试误差的闭式表达式。结果表明,在样本数量线性增长的设定下,岭回归和核回归可实现贝叶斯最优性;而当样本数量增长快于输入维度时,神经网络的表现优于这些方法,展现出在过参数化设置下的优越性。

ABSTRACT

We consider the problem of learning a target function corresponding to a deep, extensive-width, non-linear neural network with random Gaussian weights. We consider the asymptotic limit where the number of samples, the input dimension and the network width are proportionally large. We propose a closed-form expression for the Bayes-optimal test error, for regression and classification tasks. We further compute closed-form expressions for the test errors of ridge regression, kernel and random features regression. We find, in particular, that optimally regularized ridge regression, as well as kernel regression, achieve Bayes-optimal performances, while the logistic loss yields a near-optimal test error for classification. We further show numerically that when the number of samples grows faster than the dimension, ridge and kernel methods become suboptimal, while neural networks achieve test error close to zero from quadratically many samples.

研究动机与目标

  • 刻画具有高斯输入的深度广义宽度随机神经网络的信息论最小测试误差(即贝叶斯最优误差)。
  • 确定标准经验风险最小化(ERM)方法(如岭回归和核回归)是否可实现该贝叶斯最优误差。
  • 研究当样本数量增长快于输入维度时,线性方法与神经网络之间的性能差距。
  • 通过深度贝叶斯高斯等价性性质(GEP)建立理解深度学习泛化性的理论基础。

提出的方法

  • 提出深度贝叶斯高斯等价性性质(GEP)猜想,将深度随机网络在贝叶斯推断下的输出统计特性与高斯过程联系起来。
  • 利用统计力学和比例极限(n ∝ d)下的渐近分析,推导出回归与分类任务中贝叶斯最优测试误差的闭式表达式。
  • 应用复制法与洞穴法,计算岭回归、核回归及随机特征回归的渐近测试误差。
  • 通过鞍点方程与矩匹配技术,刻画高维极限下岭回归与逻辑回归的泛化误差。
  • 通过数值模拟探索 n 增长快于 d 的区域,超出 GEP 适用范围的情形。
  • 依赖于在无限宽度极限下贝叶斯推断与核方法的等价性,此处将其扩展至广义宽度网络。
Figure 1: (solid lines) Theoretical predition for the Bayes MSE ( 12 ), for a one-hidden layer rectangular neural network ( $\gamma_{1}=1$ ) with shifted ReLU activation $\sigma_{1}(\cdot)=(\cdot)_{+}-\nicefrac{{1}}{{\sqrt{2\pi}}}$ . (red crosses) Monte Carlo simulations using the Gibbs sampling alg
Figure 1: (solid lines) Theoretical predition for the Bayes MSE ( 12 ), for a one-hidden layer rectangular neural network ( $\gamma_{1}=1$ ) with shifted ReLU activation $\sigma_{1}(\cdot)=(\cdot)_{+}-\nicefrac{{1}}{{\sqrt{2\pi}}}$ . (red crosses) Monte Carlo simulations using the Gibbs sampling alg

实验结果

研究问题

  • RQ1对于具有广义宽度和高斯输入的深度随机神经网络,其渐近贝叶斯最优测试误差是什么?
  • RQ2在比例极限(n ∝ d)下,岭回归和核回归能否在该类网络中实现贝叶斯最优误差?
  • RQ3当 n 增长快于 d 时,岭回归和核方法的性能与梯度训练的神经网络相比如何?
  • RQ4深度贝叶斯高斯等价性性质(GEP)在推导闭式误差表达式中起到什么作用?
  • RQ5逻辑回归在深度随机网络的分类任务中是否能实现近似最优性能?

主要发现

  • 回归任务中,通过最优正则化处理的岭回归与核回归可实现贝叶斯最优测试误差,且推导出了最优正则化参数的闭式表达式。
  • 在分类任务中,岭回归与逻辑回归的测试误差接近但不等于贝叶斯最优误差,表明其具有近似最优性。
  • 在样本数量增长快于输入维度的区域(n ≫ d),岭回归与核方法变得次优,而梯度训练的神经网络可实现接近零的测试误差。
  • 深度贝叶斯高斯等价性性质(GEP)被猜想在比例极限下渐近成立,是推导闭式误差表达式的理论基础。
  • 数值结果证实,神经网络在过参数化设置下显著优于线性方法,凸显其归纳偏置优势。
  • 本研究确立,在高斯输入假设下,线性估计器无法从线性数量的样本中学习到目标广义宽度网络的非线性成分。
Figure 2: Targets ( 1 ) with $L=1$ (top) and $L=2$ (bottom) hidden layers , with $\sigma_{1,2}(x)=\tanh(2x)$ activation, widths $k_{1,2}=700$ , and $\Delta_{a}=\Delta_{1,2}=1,~{}\Delta=0$ , in dimension $d=500$ . The Bayes-optimal MSE ( 12 ) (dashed black) is contrasted to the replica predictions an
Figure 2: Targets ( 1 ) with $L=1$ (top) and $L=2$ (bottom) hidden layers , with $\sigma_{1,2}(x)=\tanh(2x)$ activation, widths $k_{1,2}=700$ , and $\Delta_{a}=\Delta_{1,2}=1,~{}\Delta=0$ , in dimension $d=500$ . The Bayes-optimal MSE ( 12 ) (dashed black) is contrasted to the replica predictions an

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。