Skip to main content
QUICK REVIEW

[논문 리뷰] Towards a theory of machine learning

Vitaly Vanchurin|arXiv (Cornell University)|2020. 04. 15.
Statistical Mechanics and Entropy참고 문헌 45인용 수 4
한 줄 요약

이 논문은 신경망을 정의된 상태, 가중치, 편향 및 손실 함수를 갖는 7중항으로 모델링함으로써 기계학습을 위한 통계역학 프레임워크를 제안한다. 최대 엔트로피와 분할 함수를 통해 학습의 열역학 법칙을 유도하며, 최적의 학습 효율성이 깊은 신경망에서 엔트로피와 복잡성의 상충 관계로 인해 발생함을 보여주며, 기본 물리학에서의 잠재적 시공간의 기원에 대한 함의를 제시한다.

ABSTRACT

We define a neural network as a septuple consisting of (1) a state vector, (2) an input projection, (3) an output projection, (4) a weight matrix, (5) a bias vector, (6) an activation map and (7) a loss function. We argue that the loss function can be imposed either on the boundary (i.e. input and/or output neurons) or in the bulk (i.e. hidden neurons) for both supervised and unsupervised systems. We apply the principle of maximum entropy to derive a canonical ensemble of the state vectors subject to a constraint imposed on the bulk loss function by a Lagrange multiplier (or an inverse temperature parameter). We show that in an equilibrium the canonical partition function must be a product of two factors: a function of the temperature and a function of the bias vector and weight matrix. Consequently, the total Shannon entropy consists of two terms which represent respectively a thermodynamic entropy and a complexity of the neural network. We derive the first and second laws of learning: during learning the total entropy must decrease until the system reaches an equilibrium (i.e. the second law), and the increment in the loss function must be proportional to the increment in the thermodynamic entropy plus the increment in the complexity (i.e. the first law). We calculate the entropy destruction to show that the efficiency of learning is given by the Laplacian of the total free energy which is to be maximized in an optimal neural architecture, and explain why the optimization condition is better satisfied in a deep network with a large number of hidden layers. The key properties of the model are verified numerically by training a supervised feedforward neural network using the method of stochastic gradient descent. We also discuss a possibility that the entire universe on its most fundamental level is a neural network.

연구 동기 및 목표

  • 통계역학 원리를 사용하여 지도학습 및 비지도학습을 위한 통합 이론적 프레임워크를 개발하기.
  • 높은 차원의 매개변수 공간을 가진 깊은 학습의 성공에 대한 기본적인 설명이 부족한 문제를 해결하기.
  • 경계(입력/출력 뉴런)가 아닌 부피(은닉층)에서 적용 가능한 손실 함수를 정의함으로써 비지도학습 형식을 가능하게 하기.
  • 신경망 학습 과정에 대한 평형 및 비평형 열역학 법칙을 도출하기.
  • 우주 자체가 유사한 원리에 의해 지배되는 신경망일 수 있다는 가능성 탐색하기.

제안 방법

  • 신경망을 상태 벡터, 입력/출력 투영, 가중치 행렬, 편향 벡터, 활성도 맵, 그리고 손실 함수로 구성된 7중항으로 정의한다.
  • 최대 엔트로피 원리를 적용하여, 라그랑주 승수(역온도)를 통해 부피 손실 함수로 제약된 상태 벡터의 캐노니컬 집합을 도출한다.
  • 온도에 의존하는 요소와 가중치 및 편향의 함수로 이루어진 분할 함수를 계산하여 해석적 처리를 가능하게 한다.
  • 학습의 제1법칙과 제2법칙을 도출한다: 총 엔트로피는 평형에 도달할 때까지 감소하며, 손실 증분은 열역학적 엔트로피와 복잡성 증분의 합에 비례한다.
  • 학습 효율성의 척도로 엔트로피 파괴를 도입하며, 총 자유 에너지의 라플라시안에 비례한다.
  • 유한한 활성 범위를 다루기 위해 가우시안 적분 근사와 부드러운 윈도우 함수를 사용하며, 분할 함수를 결정하는 연산자 $\hat{G}$ 를 정의한다.

실험 결과

연구 질문

  • RQ1지도학습 및 비지도학습 시스템에 대해 통합된 열역학적 기술을 어떻게 구성할 수 있는가?
  • RQ2통계역학 프레임워크 내에서 비지도학습을 가능하게 하는 부피(은닉층) 손실 함수의 역할은 무엇인가?
  • RQ3최대 엔트로피에서 도출된 캐노니컬 집합은 신경망의 평형 상태와 어떻게 관련되는가?
  • RQ4유도된 열역학 모델에 따르면, 왜 많은 은닉층을 가진 깊은 네트워크에서 학습 효율성이 더 높은가?
  • RQ5비평형 신경망 열역학으로부터 시공간과 일반 상대성 이론이 기원할 수 있는가?

주요 결과

  • 캐노니컬 분할 함수는 온도에 의존하는 항과 가중치 및 편향에 의존하는 항으로 분해되며, 총 자유 에너지가 열역학적 및 복잡성 성분으로 분리됨을 시사한다.
  • 총 샤논 엔트로피는 열역학적 엔트로피와 복잡성 항으로 분리되며, 둘 다 학습 과정에 기여한다.
  • 학습의 제1법칙은 손실의 변화가 열역학적 엔트로피 변화와 네트워크 복잡성 변화의 합에 비례한다는 것이다.
  • 학습의 제2법칙은 학습 중 총 엔트로피가 평형 상태에 도달할 때까지 감소한다는 것이다.
  • 학습 효율성이 총 자유 에너지의 라플라시안을 최대화할 때 최고조화를 이루며, 많은 은닉층을 가진 깊은 아키텍처를 선호한다.
  • 온세거 텐서에 특정 대칭성 가정을 두면 엔트로피 생성은 아인슈타인 장 방정식으로 이어지며, 일반 상대성 이론이 큰 척도에서 신경망 동역학으로부터 기원할 수 있음을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.