Skip to main content
QUICK REVIEW

[Paper Review] Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory

Paulo Tabuada, Bahman Gharesifard|arXiv (Cornell University)|Jul 12, 2020
Model Reduction and Neural Networks33 references21 citations
TL;DR

This paper establishes the universal approximation capability of deep residual neural networks by modeling them as ensemble control systems and applying geometric control theory. It proves that with specific activation functions, these networks can exactly memorize training data and approximate any continuous function on compact sets to arbitrary accuracy, providing optimal bounds on required neurons.

ABSTRACT

In this paper, we explain the universal approximation capabilities of deep residual neural networks through geometric nonlinear control. Inspired by recent work establishing links between residual networks and control systems, we provide a general sufficient condition for a residual network to have the power of universal approximation by asking the activation function, or one of its derivatives, to satisfy a quadratic differential equation. Many activation functions used in practice satisfy this assumption, exactly or approximately, and we show this property to be sufficient for an adequately deep neural network with $n+1$ neurons per layer to approximate arbitrarily well, on a compact set and with respect to the supremum norm, any continuous function from $\mathbb{R}^n$ to $\mathbb{R}^n$. We further show this result to hold for very simple architectures for which the weights only need to assume two values. The first key technical contribution consists of relating the universal approximation problem to controllability of an ensemble of control systems corresponding to a residual network and to leverage classical Lie algebraic techniques to characterize controllability. The second technical contribution is to identify monotonicity as the bridge between controllability of finite ensembles and uniform approximability on compact sets.

Motivation & Objective

  • To establish the universal approximation power of deep residual neural networks using tools from nonlinear control theory.
  • To address the limitation of prior work by focusing on deep networks with bounded width rather than unbounded width.
  • To model deep residual networks as ensemble control systems where weights serve as control inputs to steer multiple sample points.
  • To identify activation functions that ensure controllability on an open and dense submanifold of sample points.
  • To derive optimal bounds on the number of neurons required for universal approximation.

Proposed method

  • Formulate the memorization of training data as a controllability problem for an ensemble of nonlinear control systems.
  • Model each residual network layer as a control system with weights as control inputs and sample points as initial states.
  • Use Lie algebraic techniques from geometric control theory to analyze controllability of the ensemble system.
  • Apply the notion of monotonicity to extend finite-ensemble controllability to infinite-ensemble approximation.
  • Derive conditions on activation functions (specifically, those with non-vanishing higher-order derivatives) that ensure controllability.
  • Use determinant analysis of a Vandermonde-like matrix to show that controllability depends only on the highest-degree term of the activation function.

Experimental results

Research questions

  • RQ1Can deep residual neural networks with bounded width universally approximate any continuous function on compact sets?
  • RQ2What class of activation functions enables universal approximation in deep residual networks via control-theoretic methods?
  • RQ3How can the memorization of training data be framed as a controllability problem in an ensemble control system?
  • RQ4What are the optimal bounds on the number of neurons required for universal approximation in this setting?
  • RQ5How does the structure of the residual network architecture enable better approximation than shallow networks?

Key findings

  • Deep residual networks with activation functions whose second derivative is non-zero on an open and dense set can achieve universal approximation.
  • The controllability of the ensemble control system is ensured when the activation function's highest-degree term dominates, making lower-order terms irrelevant for controllability.
  • The determinant of the controllability matrix depends only on the highest-degree coefficient of the activation function, not on lower-order terms.
  • The set of weight matrices that achieve controllability is open and dense in the parameter space, implying genericity of the approximation property.
  • The number of required neurons is bounded optimally, with explicit dependence on the number of sample points and the dimension of the input space.
  • Monotonicity arguments allow the extension of finite-ensemble controllability to infinite-ensemble approximation, ensuring uniform convergence on compact sets.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.