Skip to main content
QUICK REVIEW

[论文解读] Universal Approximation Power of Deep Residual Neural Networks via Nonlinear Control Theory

Paulo Tabuada, Bahman Gharesifard|arXiv (Cornell University)|Jul 12, 2020
Model Reduction and Neural Networks参考文献 33被引用 21
一句话总结

本文通过将深度残差神经网络建模为集合控制系统并应用几何控制理论,建立了其通用逼近能力。证明了在特定激活函数下,这些网络能够精确记忆训练数据,并在紧集上以任意精度逼近任意连续函数,同时给出了所需神经元数的最优界。

ABSTRACT

In this paper, we explain the universal approximation capabilities of deep residual neural networks through geometric nonlinear control. Inspired by recent work establishing links between residual networks and control systems, we provide a general sufficient condition for a residual network to have the power of universal approximation by asking the activation function, or one of its derivatives, to satisfy a quadratic differential equation. Many activation functions used in practice satisfy this assumption, exactly or approximately, and we show this property to be sufficient for an adequately deep neural network with $n+1$ neurons per layer to approximate arbitrarily well, on a compact set and with respect to the supremum norm, any continuous function from $\mathbb{R}^n$ to $\mathbb{R}^n$. We further show this result to hold for very simple architectures for which the weights only need to assume two values. The first key technical contribution consists of relating the universal approximation problem to controllability of an ensemble of control systems corresponding to a residual network and to leverage classical Lie algebraic techniques to characterize controllability. The second technical contribution is to identify monotonicity as the bridge between controllability of finite ensembles and uniform approximability on compact sets.

研究动机与目标

  • 利用非线性控制理论工具,建立深度残差神经网络的通用逼近能力。
  • 通过聚焦于宽度有界的深层网络而非宽度无界的网络,克服先前研究的局限性。
  • 将深度残差网络建模为集合控制系统,其中权重作为控制输入,用于引导多个样本点。
  • 识别出能确保在样本点的开且稠密子流形上可实现可控性的激活函数。
  • 推导出实现通用逼近所需神经元数的最优界。

提出的方法

  • 将训练数据的记忆化问题表述为一组非线性控制系统的可控性问题。
  • 将每个残差网络层建模为一个控制系统,其中权重作为控制输入,样本点作为初始状态。
  • 利用几何控制理论中的李代数技术分析集合系统的可控性。
  • 应用单调性概念,将有限集合的可控性扩展至无限集合的逼近。
  • 推导出激活函数的条件(特别是高阶导数非零的函数),以确保可控性。
  • 通过分析类似范德蒙德矩阵的可控性矩阵的行列式,证明可控性仅取决于激活函数的最高次项。

实验结果

研究问题

  • RQ1具有有界宽度的深度残差神经网络是否能对紧集上的任意连续函数实现通用逼近?
  • RQ2通过控制理论方法,哪类激活函数可实现深度残差网络的通用逼近?
  • RQ3如何将训练数据的记忆化问题框架化为集合控制系统的可控性问题?
  • RQ4在此设置下,实现通用逼近所需的神经元数的最优界是什么?
  • RQ5残差网络架构的结构如何使其逼近能力优于浅层网络?

主要发现

  • 当激活函数的二阶导数在开且稠密集上非零时,具有此类激活函数的深度残差网络可实现通用逼近。
  • 当激活函数的最高次项占主导地位时,集合控制系统的可控性得以保证,此时低阶项对可控性无影响。
  • 可控性矩阵的行列式仅依赖于激活函数的最高次项系数,而不依赖于低阶项。
  • 实现可控性的权重矩阵集合在参数空间中为开且稠密,表明逼近性质具有普遍性。
  • 所需神经元数的界限是最优的,且显式依赖于样本点数量和输入空间维度。
  • 通过单调性论证,可将有限集合的可控性扩展至无限集合逼近,确保在紧集上的一致收敛。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。