Skip to main content
QUICK REVIEW

[논문 리뷰] Universal Approximation in Dropout Neural Networks

Oxana A. Manita, Mark A. Peletier|arXiv (Cornell University)|2020. 12. 18.
Neural Networks and Applications인용 수 11
한 줄 요약

이 논문은 드롭아웃 신경망에 대한 일반적 근사 정리들을 수립하여, 드롭아웃 네트워크의 랜덤 모드와 결정론적 모드가 임의의 가측 함수를 임의의 정밀도로 근사할 수 있음을 증명한다. 기대값의 대수적 성질과 재귀적 네트워크 전개를 활용함으로써, 활성화 함수와 필터 분포에 대한 최소한의 가정 하에 드롭아웃 네트워크가 입력층 간선이 랜덤하게 제거되는 상황에서도 일반적 근사성을 유지함을 보여준다.

ABSTRACT

We prove two universal approximation theorems for a range of dropout neural networks. These are feed-forward neural networks in which each edge is given a random $\{0,1\}$-valued filter, that have two modes of operation: in the first each edge output is multiplied by its random filter, resulting in a random output, while in the second each edge output is multiplied by the expectation of its filter, leading to a deterministic output. It is common to use the random mode during training and the deterministic mode during testing and prediction. Both theorems are of the following form: Given a function to approximate and a threshold $\varepsilon>0$, there exists a dropout network that is $\varepsilon$-close in probability and in $L^q$. The first theorem applies to dropout networks in the random mode. It assumes little on the activation function, applies to a wide class of networks, and can even be applied to approximation schemes other than neural networks. The core is an algebraic property that shows that deterministic networks can be exactly matched in expectation by random networks. The second theorem makes stronger assumptions and gives a stronger result. Given a function to approximate, it provides existence of a network that approximates in both modes simultaneously. Proof components are a recursive replacement of edges by independent copies, and a special first-layer replacement that couples the resulting larger network to the input. The functions to be approximated are assumed to be elements of general normed spaces, and the approximations are measured in the corresponding norms. The networks are constructed explicitly. Because of the different methods of proof, the two results give independent insight into the approximation properties of random dropout networks. With this, we establish that dropout neural networks broadly satisfy a universal-approximation property.

연구 동기 및 목표

  • 드롭아웃 신경망이 랜덤한 간선 필터링에도 불구하고 일반적 근사 성질을 만족함을 수립하는 것.
  • 드롭아웃의 확률적 성격이 임의의 함수 근사 능력을 방해하는지 분석하는 것.
  • 랜덤 모드와 결정론적 모드의 드롭아웃 네트워크가 모두 임의의 정밀도로 근사 정확도를 달성할 수 있음을 보여주는 것.
  • ReLU 네트워크를 초월해 광범위한 활성화 함수와 필터 분포에 대한 일반적 근사 결과를 확장하는 것.
  • Lq 및 확률 노름에서 목표 함수를 근사하는 드롭아웃 네트워크의 명시적 구성 제공

제안 방법

  • 기대값에서 결정론적 네트워크와 일치하는 대수적 성질을 활용하여 랜덤 모드의 드롭아웃 네트워크에 대한 일반적 근사 정리를 증명한다.
  • 간선을 독립적인 복제본으로 대체하는 재귀적 구성 방법을 통해 분산을 제어하고 확률 수렴을 보장한다.
  • 확장된 네트워크를 입력에 연결하는 전용 첫 번째 레이어 대체를 도입하여 출력 분포를 제어한다.
  • 기대값–분산 분해를 사용하여 Lq 및 확률 노름에서의 근사 오차를 경계한다.
  • 일반적인 노름이 부여된 함수 공간에 결과를 적용하여 근사 오차가 적절한 준노름에서 측정됨을 보장한다.
  • 네트워크 출력의 순열 대칭성을 활용하여 입력층 드롭아웃이 일반적 근사성을 방해하지 않음을 보여준다.

실험 결과

연구 질문

  • RQ1드롭아웃 신경망은 랜덤 모드와 결정론적 모드에서 임의의 가측 함수를 임의의 정밀도로 근사할 수 있는가?
  • RQ2드롭아웃에 의해 도입된 랜덤성은 표준 신경망의 일반적 근사 성질을 약화시키는가?
  • RQ3활성화 함수의 성질과 필터 분포에 대한 가정이 드롭아웃 네트워크의 근사 능력에 어느 정도 영향을 미치는가?
  • RQ4입력층 간선이 랜덤하게 제거되는 상황에서도 일반적 근사가 달성될 수 있는가?
  • RQ5추론 시 기대값 대체 드롭아웃이 강력한 성능을 보이는 이론적 근거는 무엇인가?

주요 결과

  • 랜덤 모드에서 드롭아웃 네트워크에 일반적 근사 정리가 성립한다: 임의의 ε > 0 과 목표 함수 ζ에 대해, 근사 오차가 ε를 초과할 확률이 ε 미만이 되는 드롭아웃 네트워크가 존재한다.
  • 근사 오차는 Lq 노름에서 경계된다: 임의의 ε > 0 에 대해, 오차의 Lq 노름이 ε 이하가 되는 드롭아웃 네트워크가 존재한다.
  • 첫 번째 정리는 엣지 기반 드롭컨넥트와 노드 기반 드롭아웃을 포함한 광범위한 활성화 함수 및 필터 분포 클래스에 적용되며, 최소한의 가정으로도 성립한다.
  • 두 번째 정리는 랜덤 모드와 결정론적 모드에서 동시에 목표 함수를 잘 근사하는 단일 네트워크를 구성한다.
  • 증명은 첫 번째 레이어를 병렬로 복제함으로써 확장하고, 순열 대칭성을 활용하여 출력 분포를 제어하는 데 의존한다.
  • 입력층 드롭아웃이 존재하더라도, 랜덤 실현치들이 서로 순열 치환임의 높은 확률을 가짐으로써 일반적 근사성이 유지된다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.