Skip to main content
QUICK REVIEW

[论文解读] Sharp Bounds on the Approximation Rates, Metric Entropy, and $n$-widths of Shallow Neural Networks

Jonathan W. Siegel, Jinchao Xu|arXiv (Cornell University)|Jan 29, 2021
Advanced Numerical Analysis Techniques参考文献 68被引用 9
一句话总结

该论文通过引入一个平滑参数化的字典框架,建立了浅层神经网络在 ReLU^k 和 Sigmoid 激活函数下的逼近速率、度量熵和 n-宽度的精确界限。论文证明了变分空间上的最优收敛速率,表明 ReLU^k 网络可实现阶为 $ n^{-1/2 - 1/(2d)} $ 的最佳可能逼近误差衰减率,该速率对给定函数类是紧致的。

ABSTRACT

In this article, we study approximation properties of the variation spaces corresponding to shallow neural networks with a variety of activation functions. We introduce two main tools for estimating the metric entropy, approximation rates, and $n$-widths of these spaces. First, we introduce the notion of a smoothly parameterized dictionary and give upper bounds on the non-linear approximation rates, metric entropy and $n$-widths of their absolute convex hull. The upper bounds depend upon the order of smoothness of the parameterization. This result is applied to dictionaries of ridge functions corresponding to shallow neural networks, and they improve upon existing results in many cases. Next, we provide a method for lower bounding the metric entropy and $n$-widths of variation spaces which contain certain classes of ridge functions. This result gives sharp lower bounds on the $L^2$-approximation rates, metric entropy, and $n$-widths for variation spaces corresponding to neural networks with a range of important activation functions, including ReLU$^k$ activation functions and sigmoidal activation functions with bounded variation.

研究动机与目标

  • 建立浅层神经网络在各种激活函数下的逼近速率、度量熵和 n-宽度的精确上下界。
  • 基于平滑参数化字典,构建一个分析巴拿赫空间中非线性字典逼近的理论框架。
  • 刻画浅层网络在与脊函数相关的变分空间上的最优逼近性能。
  • 证明 ReLU^k 网络在其变分空间中实现了最佳可能的收敛速率,确认了现有结果的最优性。

提出的方法

  • 在巴拿赫空间中引入平滑参数化字典的概念,其中字典元素由权重和偏置的光滑函数参数化。
  • 基于参数化平滑度阶数,推导出此类字典的绝对凸包的逼近速率、度量熵和 n-宽度的上界。
  • 将上界框架应用于由浅层神经网络中 ReLU^k 和 Sigmoid 激活函数产生的脊函数字典。
  • 通过构造具有受控 $ L^2 $-距离的函数网并结合几何测度论论证,提出一种新颖的度量熵和 n-宽度下界估计方法。
  • 利用下界技术证明,ReLU^k 网络的逼近速率 $ n^{-1/2 - 1/(2d)} $ 是最优的,且对具有有界变差的 Sigmoid 网络也成立。
  • 建立 ReLU^k 网络的变分空间与参数化字典的凸包之间的等价关系,从而实现对熵和 n-宽度的精确估计。

实验结果

研究问题

  • RQ1浅层 ReLU^k 和 Sigmoid 激活函数神经网络的逼近速率的精确上下界是什么?
  • RQ2浅层网络变分空间的度量熵和 n-宽度如何随网络宽度和维度变化?
  • RQ3能否证明 $ n^{-1/2 - 1/(2d)} $ 的逼近速率在 $ L^2 $-范数下对 ReLU^k 网络是最优的?
  • RQ4字典参数化的平滑度如何影响逼近和熵的界?
  • RQ5这些界在多大程度上依赖于激活函数的选择,特别是对于非有界函数如 ReLU^k?

主要发现

  • 浅层 ReLU^k 网络在 $ L^2 $ 范数下的逼近速率上下界均为 $ O(n^{-1/2 - 1/(2d)}) $,证明了该速率的最优性。
  • 对于具有有界变差的 Sigmoid 激活函数,相同的逼近速率 $ n^{-1/2 - 1/(2d)} $ 被证明是最优的,即使系数在 $ \ell^\infty $ 中有界。
  • 变分空间 $ B_1(\mathbb{P}_k^d) $ 的度量熵满足 $ \epsilon_n(B_1(\mathbb{P}_k^d)) \gtrsim n^{-1/2 - 1/(2d)} $,与上界一致,确认了紧致性。
  • 变分空间 $ B_1(\mathbb{P}_k^d) $ 的 $ n $-宽度以速率 $ n^{-1/2 - 1/(2d)} $ 衰减,这是该类别的最优速率。
  • 通过构造具有受控 $ L^2 $-距离的脊函数网,并结合在非零差集上的测度论论证,可获得精确的下界。
  • 结果可推广至具有有界变差的一般 Sigmoid 函数,表明在有界变差之外更强的光滑性假设并不能改善逼近速率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。