Skip to main content
QUICK REVIEW

[Paper Review] Nonlinear Approximation and (Deep) ReLU Networks

Ingrid Daubechies, Ronald DeVore|arXiv (Cornell University)|May 5, 2019
Neural Networks and Applications6 references104 citations
TL;DR

The paper analyzes the expressive power of deep ReLU networks for univariate functions, showing depth yields approximation benefits beyond free-knot linear splines and introducing special network constructions to prove containment and compositional capabilities.

ABSTRACT

This article is concerned with the approximation and expressive powers of deep neural networks. This is an active research area currently producing many interesting papers. The results most commonly found in the literature prove that neural networks approximate functions with classical smoothness to the same accuracy as classical linear methods of approximation, e.g. approximation by polynomials or by piecewise polynomials on prescribed partitions. However, approximation by neural networks depending on n parameters is a form of nonlinear approximation and as such should be compared with other nonlinear methods such as variable knot splines or n-term approximation from dictionaries. The performance of neural networks in targeted applications such as machine learning indicate that they actually possess even greater approximation power than these traditional methods of nonlinear approximation. The main results of this article prove that this is indeed the case. This is done by exhibiting large classes of functions which can be efficiently captured by neural networks where classical nonlinear methods fall short of the task. The present article purposefully limits itself to studying the approximation of univariate functions by ReLU networks. Many generalizations to functions of several variables and other activation functions can be envisioned. However, even in this simplest of settings considered here, a theory that completely quantifies the approximation power of neural networks is still lacking.

Motivation & Objective

  • Assess the approximation capabilities of deep ReLU networks for univariate functions compared to classical nonlinear methods.
  • Establish that fixed-width deep ReLU networks with comparable parameter counts can match free knot linear splines and surpass them in expressive power.
  • Demonstrate how depth enables efficient representation through composition and self-similar structures.
  • Introduce and analyze special network constructions that preserve or enhance approximation power.
  • Discuss implications for data fitting bias and potential practical architectures.

Proposed method

  • Define the function classes realized by ReLU networks with width W and depth L, denoted Upsilon^{W,L}.
  • Compare Upsilon^{W,L} to the nonlinear spline class Sigma_n of CPwL functions with n breakpoints.
  • Prove Sigma_n is contained in Upsilon^{W,L} with n(W,L) comparable to n, using a two-layer special network construction.
  • Develop a two-layer construction that generates CPwL functions with controlled breakpoints via hat functions and principal breakpoints.
  • Extend to compositions and sums of CPwL components and analyze how depth contributes to expressive power.
  • Provide theoretical results on composition and sum stability within special networks and standard networks.

Experimental results

Research questions

  • RQ1Can fixed-width deep ReLU networks approximate free knot linear splines with a comparable parameter budget?
  • RQ2Do deep compositions of CPwL functions yield expressive power beyond what Sigma_n offers?
  • RQ3How does depth interact with width to influence the approximation capacity relative to classical nonlinear methods?
  • RQ4What network architectures (special networks) can demonstrate constructive containment and additional expressive power?
  • RQ5What are the implications for bias in data fitting when using deep ReLU networks?

Key findings

  • Fixed-width ReLU networks with n parameters can approximate Sigma_n functions up to a constant factor in n(W,L).
  • Sigma_n is contained in Upsilon^{W,L} with depth and width enabling roughly n breakpoints to be represented.
  • Deep networks can generate functions with many breakpoints via composition, surpassing polynomial growth tied to n.
  • Special network constructions (SC and CC channels) provide a framework to realize CPwL functions and their sums/compositions.
  • Composition of CPwL functions can be realized with bounded width and depth while controlling parameter counts.
  • The paper identifies classes of self-similar and trigonometric-like functions that deep ReLU networks can emulate efficiently.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.