[论文解读] High-Order Approximation Rates for Shallow Neural Networks with Cosine and ReLU$^k$ Activation Functions
本文建立了使用余弦和ReLU$^k$激活函数的浅层神经网络的最优高阶逼近率。证明了ReLU$^k$网络在维度$d$无关的情况下,可实现$O(n^{-(k+1)}\log n)$的$L^2$逼近率,优于经典方法(后者需随维度增加平滑性或多项式次数)。该结果适用于谱Barron空间中具有平滑性指标$s$的函数,且通过匹配下界证明了逼近率的最优性。
We study the approximation properties of shallow neural networks with an activation function which is a power of the rectified linear unit. Specifically, we consider the dependence of the approximation rate on the dimension and the smoothness in the spectral Barron space of the underlying function $f$ to be approximated. We show that as the smoothness index $s$ of $f$ increases, shallow neural networks with ReLU$^k$ activation function obtain an improved approximation rate up to a best possible rate of $O(n^{-(k+1)}\log(n))$ in $L^2$, independent of the dimension $d$. The significance of this result is that the activation function ReLU$^k$ is fixed independent of the dimension, while for classical methods the degree of polynomial approximation or the smoothness of the wavelets used would have to increase in order to take advantage of the dimension dependent smoothness of $f$. In addition, we derive improved approximation rates for shallow neural networks with cosine activation function on the spectral Barron space. Finally, we prove lower bounds showing that the approximation rates attained are optimal under the given assumptions.
研究动机与目标
- 分析浅层神经网络在谱Barron空间中使用ReLU$^k$与余弦激活函数时的逼近能力。
- 量化目标函数$f$的平滑性$s$增加时,逼近率的改善程度。
- 为ReLU$^k$网络建立维度无关的逼近率,克服经典方法需随维度增加平滑性的局限。
- 推导匹配的下界,证明所获逼近率的最优性。
提出的方法
- 分析在谱Barron空间$\mathcal{B}^s(\Omega)$中进行,该空间通过平滑性指标$s$控制傅里叶域中的衰减。
- 对于ReLU$^k$网络,方法通过构造形如$\sigma(\omega_i \cdot x + b_i)$的径向基函数线性组合实现逼近,其中$\sigma(x) = [\max(0,x)]^k$。
- 证明结合了傅里叶分析、覆盖论证与熵估计,以界定在给定误差范围内逼近所需的不同网络配置数量。
- 通过将问题约化为一维情形,并利用解析函数分段多项式逼近的已知结果,推导出下界。
- 余弦激活的论证采用类似的熵与覆盖技术,通过扰动傅里叶系数与频率,构造出一组可区分的函数。
- 基于熵与覆盖数的矛盾论证用于证明最优性,表明在给定假设下,不可能存在更快的逼近率。
实验结果
研究问题
- RQ1浅层ReLU$^k$激活神经网络能否实现不依赖输入维度$d$的、随目标函数平滑性$s$提升而改善的逼近率?
- RQ2在谱Barron空间$\mathcal{B}^s(\Omega)$中,浅层ReLU$^k$网络在$L^2$范数下可实现的最佳逼近率是多少?
- RQ3在相同平滑性假设下,余弦激活网络的逼近率与ReLU$^k$网络相比如何?
- RQ4所推导的逼近率是否最优,或可通过其他构造进一步改进?
主要发现
- 浅层ReLU$^k$网络在谱Barron空间$\mathcal{B}^s(\Omega)$中,对任意输入维度$d$,可实现$O(n^{-(k+1)}\log n)$的$L^2$逼近率。
- 该速率是最优的,由匹配下界证明:对任意$m \leq s$,存在$f \in \mathcal{B}^s(\Omega)$,使得$\|f - f_n\|_{H^m(\Omega)} \gtrsim n^{m-(k+1)}$。
- ReLU$^k$网络的逼近率随$k$增加而改善,最优速率可达$O(n^{-(k+1)}\log n)$,无法进一步提升。
- 对于余弦激活,推导出谱Barron空间中的改进逼近率,尽管在摘要中未以闭式表达给出确切速率。
- 结果表明,ReLU$^k$网络可实现维度无关的逼近速率,而经典方法(如有限元或小波)则需随$d$增加平滑性或次数。
- 通过基于熵与覆盖数的矛盾论证,证明了速率的最优性,表明所构造的函数集合无法被少于$\sim R^{(2s + \frac{5}{2}d - t)dR^{d-t}}$个元素覆盖。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。