[论文解读] Approximation Rates for Neural Networks with General Activation Functions
本文通过利用傅里叶分析和Barron范数约束,建立了具有广义激活函数的两层神经网络的维数无关逼近率。对于多项式衰减的激活函数,证明了 $ n^{-1/2} $ 的收敛速率;对于有界可积的激活函数,收敛速率较弱,为 $ n^{-1/4} $,并在温和条件下通过分层采样进一步提升了逼近速率。
We prove some new results concerning the approximation rate of neural networks with general activation functions. Our first result concerns the rate of approximation of a two layer neural network with a polynomially-decaying non-sigmoidal activation function. We extend the dimension independent approximation rates previously obtained to this new class of activation functions. Our second result gives a weaker, but still dimension independent, approximation rate for a larger class of activation functions, removing the polynomial decay assumption. This result applies to any bounded, integrable activation function. Finally, we show that a stratified sampling approach can be used to improve the approximation rate for polynomially decaying activation functions under mild additional assumptions.
研究动机与目标
- 将两层神经网络的维数无关逼近率从Sigmoid型或周期型激活函数扩展至更广义的激活函数。
- 利用傅里叶分析与Barron范数约束,为具有多项式衰减特性的激活函数建立更优的收敛速率。
- 将结果推广至更广泛的有界可积激活函数类,且无需依赖多项式衰减假设。
- 在温和的附加假设下,通过分层采样技术提升逼近速率。
- 为浅层网络如何通过Barron范数正则性克服维数灾难提供理论基础。
提出的方法
- 利用傅里叶分析,将逼近误差表示为目标函数与激活函数傅里叶变换的函数形式。
- 采用Barron范数 $ \|f\|_{\mathcal{B}^{s}} $ 作为光滑性度量,以在Sobolev空间中界定逼近误差。
- 应用激活函数的近似积分表示,推导出广义激活函数的维数无关上界。
- 引入一种分层采样策略,将参数空间划分为若干子区域,通过降低逼近误差的方差来提升收敛速率。
- 基于随机参数的期望,推导出在 $ H^m(\Omega) $-范数下的误差界,并通过优化采样参数最小化误差。
- 采用 $ L^2 $-范数误差估计与浓度不等式,获得高概率逼近保证。
实验结果
研究问题
- RQ1能否将两层神经网络在多项式衰减激活函数下的 $ n^{-1/2} $ 逼近速率推广至更广义的激活函数?
- RQ2对于无多项式衰减的广义有界可积激活函数,最佳可能的逼近速率是什么?
- RQ3在温和的正则性假设下,分层采样能否提升神经网络逼近的收敛速率?
- RQ4Barron范数如何影响浅层网络在高维空间中的逼近能力?
- RQ5对于广义激活函数,是否可能实现与ReLU或Sigmoid函数相同的 $ n^{-1/2} $ 收敛速率?
主要发现
- 对于多项式衰减的激活函数,本文建立了维数无关的逼近速率 $ n^{-1/2} $,相较于先前结果去除了对数因子,实现改进。
- 对于广义有界可积激活函数(如有界变差函数),本文证明了较弱但同样维数无关的速率 $ n^{-1/4} $,将结果拓展至Sigmoid型或周期型函数之外。
- 通过应用分层采样,本文在最优参数选择下,为多项式衰减激活函数获得了改进的收敛速率 $ n^{-1/2 - \frac{\min(p-1,\epsilon)}{(d+1)(2 + \min(p-1,\epsilon))}} $。
- 对于周期型激活函数,分层采样方法实现了 $ n^{-1/2 - \frac{\epsilon}{(d+1)(2 + \epsilon)}} $ 的速率,优于基线的 $ n^{-1/2} $ 速率。
- 所有结果均基于目标函数具有有限Barron范数 $ \|f\|_{\mathcal{B}^{m+\epsilon}} $ 的假设,该光滑性条件弱于以往基于Sobolev空间的分析所采用的条件。
- 理论框架表明,当目标函数具有有界Barron范数时,浅层网络即使在缺乏强光滑性假设下,也能克服维数灾难。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。