Skip to main content
QUICK REVIEW

[论文解读] Deep Nonparametric Regression on Approximate Manifolds: Non-Asymptotic Error Bounds with Polynomial Prefactors

Yuling Jiao, Guohao Shen|arXiv (Cornell University)|Apr 14, 2021
Topological and Geometric Data Analysis参考文献 94被引用 11
一句话总结

本文在低维流形假设下建立了深度神经网络回归的非渐近误差界,表明预测误差的收敛速率达到极小极大最优,且对维度的依赖为多项式而非指数形式。本文提出了ReLU网络的新近似误差界,并定义了网络相对效率以比较不同网络结构,证明当数据位于或接近低维流形时,深度网络可克服维度灾难。

ABSTRACT

We study the properties of nonparametric least squares regression using deep neural networks. We derive non-asymptotic upper bounds for the prediction error of the empirical risk minimizer of feedforward deep neural regression. Our error bounds achieve minimax optimal rate and significantly improve over the existing ones in the sense that they depend polynomially on the dimension of the predictor, instead of exponentially on dimension. We show that the neural regression estimator can circumvent the curse of dimensionality under the assumption that the predictor is supported on an approximate low-dimensional manifold or a set with low Minkowski dimension. We also establish the optimal convergence rate under the exact manifold support assumption. We investigate how the prediction error of the neural regression estimator depends on the structure of neural networks and propose a notion of network relative efficiency between two types of neural networks, which provides a quantitative measure for evaluating the relative merits of different network structures. To establish these results, we derive a novel approximation error bound for the Hölder smooth functions with a positive smoothness index using ReLU activated neural networks, which may be of independent interest. Our results are derived under weaker assumptions on the data distribution and the neural network structure than those in the existing literature.

研究动机与目标

  • 建立深度神经网络回归的非渐近预测误差界,使其在输入维度上的依赖为多项式,避免指数依赖。
  • 证明当预测变量支持在低维流形或具有低Minkowski维数的集合上时,深度神经网络可规避维度灾难。
  • 利用ReLU激活的前馈神经网络,为Hölder光滑函数推导新的近似误差界。
  • 引入并形式化网络相对效率的概念,以定量比较不同神经网络结构。
  • 相比先前工作,放宽对数据分布和网络结构的假设,提升方法的适用范围。

提出的方法

  • 通过ReLU网络推导深度非参数回归中经验风险最小化器的非渐近上界。
  • 为使用受控宽度与深度的ReLU激活前馈神经网络,建立Hölder光滑函数的新近似误差界。
  • 利用随机投影将高维数据嵌入低维空间,同时保持成对距离在误差因子δ以内。
  • 通过ε-熵的Johnson-Lindenstrauss型集中不等式,以内在维数d*控制嵌入维数d₀。
  • 结合近似误差与估计误差项,推导出具有维度多项式系数的复合风险界。
  • 引入网络相对效率的概念,以在相同参数预算下比较不同网络结构的收敛速率。

实验结果

研究问题

  • RQ1在低维流形假设下,深度神经网络回归能否实现收敛速率的极小极大最优,且对维度的依赖为多项式而非指数?
  • RQ2在高维情形下,ReLU网络对Hölder光滑函数的近似误差如何随维度变化?
  • RQ3网络结构(宽度、深度、参数量)与非参数回归中预测误差之间的关系为何?
  • RQ4如何定量比较不同神经网络结构在收敛速率方面的效率?
  • RQ5在何种条件下,深度网络可在非参数回归中克服维度灾难?

主要发现

  • 深度神经网络估计器的预测误差被一个随n^{-2β/(2β + d₀)}变化的项所界定,其中d₀为内在维数,该结果在低维流形支持下达到极小极大最优速率。
  • 误差界对环境维数d表现出多项式依赖,具体为d^{1/2}d₀^{⌊β⌋ + (β∨1+1)/2},而非指数依赖,从而缓解了维度灾难。
  • 推导出Hölder光滑函数的新近似误差界,表明ReLU网络可使误差以(NM)^{-2β/d₀}的速率衰减,其中N×M为网络规模。
  • 所提出的网络相对效率度量可实现对不同网络类型的定量比较,表明在相同参数预算下,更深或更宽的架构可能更高效。
  • 结果在数据分布和网络结构假设上弱于先前工作,增强了其实际适用性。
  • 理论框架通过随机投影理论与基于ε-熵的集中不等式联合验证,确保在高维设置下的稳健性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。