[Paper Review] Optimal rates of approximation by shallow ReLU$^k$ neural networks and applications to nonparametric regression
This paper establishes optimal approximation rates for shallow ReLU$^k$ neural networks by analyzing variation norms of functions in the unit ball of Hölder spaces. It proves that shallow networks achieve minimax optimal convergence rates for nonparametric regression of Hölder-smooth functions, and extends these results to over-parameterized and convolutional neural networks with nearly optimal rates.
We study the approximation capacity of some variation spaces corresponding to shallow ReLU$^k$ neural networks. It is shown that sufficiently smooth functions are contained in these spaces with finite variation norms. For functions with less smoothness, the approximation rates in terms of the variation norm are established. Using these results, we are able to prove the optimal approximation rates in terms of the number of neurons for shallow ReLU$^k$ neural networks. It is also shown how these results can be used to derive approximation bounds for deep neural networks and convolutional neural networks (CNNs). As applications, we study convergence rates for nonparametric regression using three ReLU neural network models: shallow neural network, over-parameterized neural network, and CNN. In particular, we show that shallow neural networks can achieve the minimax optimal rates for learning Hölder functions, which complements recent results for deep neural networks. It is also proven that over-parameterized (deep or shallow) neural networks can achieve nearly optimal rates for nonparametric regression.
Motivation & Objective
- To establish optimal approximation rates for shallow ReLU$^k$ neural networks in terms of the number of neurons.
- To analyze the approximation capacity of variation spaces associated with shallow ReLU$^k$ networks.
- To derive convergence rates for nonparametric regression using three neural network models: shallow, over-parameterized, and convolutional neural networks.
- To show that shallow networks achieve minimax optimal rates for Hölder-smooth functions, complementing existing deep network results.
- To extend the analysis to deep and convolutional networks using the derived approximation bounds.
Proposed method
- Define the function class $\mathcal{F}_{\sigma_k}(M)$ as the set of infinite-width shallow ReLU$^k$ networks with bounded total variation of the measure $\mu$, i.e., $\|\mu\| \leq M$.
- Use the integral representation of functions via spherical harmonics and the Maurey-type random sampling argument to bound approximation error in terms of variation norms.
- Establish the approximation rate $\|h - f\|_{L^\infty} \lesssim M^{-2\alpha/(d+2k+1-2\alpha)}$ for $h \in \mathcal{H}^\alpha$ when $\alpha < (d+2k+1)/2$.
- Combine the variation norm approximation result with random approximation bounds from [55, 57] to derive optimal rates for finite-width shallow networks.
- Apply the approximation bounds to derive generalization error bounds for nonparametric regression via empirical risk minimization with truncation.
- Use pseudo-dimension and covering number arguments to control the complexity of the network classes, especially for CNNs and over-parameterized models.
Experimental results
Research questions
- RQ1What is the optimal approximation rate achievable by shallow ReLU$^k$ neural networks for functions in the Hölder space $\mathcal{H}^\alpha$?
- RQ2Can shallow ReLU$^k$ networks achieve minimax optimal convergence rates in nonparametric regression for Hölder-smooth functions?
- RQ3How do the approximation rates of shallow networks extend to deep and convolutional neural networks?
- RQ4Can over-parameterized neural networks achieve nearly optimal rates for nonparametric regression?
- RQ5What role does the activation function's power $k$ play in determining the approximation rate and the smoothness threshold?
Key findings
- For $\alpha < (d+2k+1)/2$, the approximation error of shallow ReLU$^k$ networks in the $L^\infty$ norm decays as $\mathcal{O}(M^{-2\alpha/(d+2k+1-2\alpha)})$, where $M$ is the total variation of the measure.
- Shallow ReLU$^k$ networks achieve the optimal approximation rate $\mathcal{O}(N^{-\alpha/d})$ for $\mathcal{H}^\alpha$ when $\alpha < (d+2k+1)/2$, generalizing Mhaskar's result to ReLU$^k$.
- Shallow neural networks achieve minimax optimal convergence rates for learning Hölder functions, matching the best possible rate in nonparametric regression.
- Over-parameterized (deep or shallow) neural networks achieve nearly optimal convergence rates for nonparametric regression, with error bounds of order $\mathcal{O}(n^{-\alpha/(d+\alpha)}(\log n)^4)$.
- Convolutional neural networks with ReLU activation ($k=1$) achieve convergence rates of $\mathcal{O}(n^{-(d+3)/(3d+3)}(\log n)^4)$ for $\mathcal{F}_\sigma(1)$, which is slightly better than existing results.
- The approximation results extend to Sobolev norms, suggesting potential for generalization to other smoothness spaces, though this is left for future work.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.