Skip to main content
QUICK REVIEW

[논문 리뷰] Gaussian Process Behaviour in Wide Deep Neural Networks

Alexander Matthews, Mark Rowland|arXiv (Cornell University)|2018. 04. 30.
Gaussian Processes and Bayesian Inference인용 수 90
한 줄 요약

이 논문은 깊고 넓은, 다중 히든 레이어를 가진 완전 연결 신경망이 폭이 커지면 분포상 가우시안 프로세스로 수렴한다는 것을 mild 조건하에 보이고, MMD를 이용해 Gaussian process 아날로그 및 정확한 베이지안 신경망과의 비교로 경험적으로 검증한다.

ABSTRACT

Whilst deep neural networks have shown great empirical success, there is still much work to be done to understand their theoretical properties. In this paper, we study the relationship between random, wide, fully connected, feedforward networks with more than one hidden layer and Gaussian processes with a recursive kernel definition. We show that, under broad conditions, as we make the architecture increasingly wide, the implied random function converges in distribution to a Gaussian process, formalising and extending existing results by Neal (1996) to deep networks. To evaluate convergence rates empirically, we use maximum mean discrepancy. We then compare finite Bayesian deep networks from the literature to Gaussian processes in terms of the key predictive quantities of interest, finding that in some cases the agreement can be very close. We discuss the desirability of Gaussian process behaviour and review non-Gaussian alternative models from the literature.

연구 동기 및 목표

  • 단순 다중 히든 레이어를 가진 무작위 완전 연결 네트워크에 대한 이론적 이해를 확장한다.
  • 와이드 deep 네트워크가 broad condition 하에서 Gaussian processes로 수렴함을 증명한다.
  • 최대 평균 차이(MMD)를 사용한 수렴 속도에 대한 실증 평가를 수행한다.
  • 제한된 Bayesian 딥 네트워크를 예측적 양에서 Gaussian processes와 비교한다.
  • Bayesian 딥 러닝 및 초기화/다이나믹스에 대한 시사점을 논의한다.

제안 방법

  • 가중치와 바이어스에 대해 표준 무작위 정규 사전분포를 갖는 D개의 히든 레이어를 가진 완전 연결 네트워크를 정의한다.
  • Variance explosion을 피하기 위해 Neal(1996)에 따라 너비에 따라 가중치 분산을 스케일링한다.
  • 다변량 중심극한정리를 사용하여 층 활성화의 결합 분포가 다변량 정규분포로 수렴함을 보이고, 극한에서 Gaussian process를 유도한다(정리 4).
  • 비선형성 φ의 선형 포괄 속성(|φ(u)| ≤ c + m|u|)을 가정한다.
  • 레이어 간 한정 공분산 구조를 특징짓기 위한 재귀 보조정리(Lemma 2)를 사용한다.
  • 유한한 네트워크와 GP 아날로그 간의 MMD를 통해 GP로의 수렴을 측정한다.

실험 결과

연구 질문

  • RQ1폭이 증가함에 따라 깊고 넓은 신경망이 분포적으로 Gaussian process로 수렴하는 조건은 무엇인가?
  • RQ2레이어당 증가하는 너비 증가 방식이 GP로의 수렴에 어떤 영향을 미치는가?
  • RQ3깊이와 너비에 따른 Gaussian process로의 수렴 속도는 어떻게 되는가?
  • RQ4일반 데이터셋과 사전분포에서 유한한 Bayesian 딥 네트워크가 GP 예측과 얼마나 잘 일치하는가?

주요 결과

  • Rigorous result (Theorem 4) shows convergence to a Gaussian process for any fixed number of hidden layers with strictly increasing width functions.
  • The limiting GP has zero mean and a covariance determined by a recursion (Lemma 2).
  • Empirical MMD experiments show finite networks increasingly resemble their GP analogues as width grows, with slower convergence for deeper nets.
  • Among six datasets, five show close agreement between exact GP inference and finite Bayesian neural networks using MCMC.
  • Different width growth schemes (identity, largest last, largest first) still lead to GP convergence as width increases, confirming independence from the specific width function shape under the theorem.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.