Skip to main content
QUICK REVIEW

[논문 리뷰] Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models

Benjamin Hoover, Hendrik Strobelt|arXiv (Cornell University)|2023. 09. 28.
Neural Networks and ApplicationsComputer Science인용 수 3
한 줄 요약

이 논문은 확산 모델과 연상 기억 사이의 깊은 수학적 동치성을 드러내며, 둘 다 고정점 기준의 안정점으로 향하는 리아푸노프 에너지 함수에 대한 경사 하강법으로 작동한다는 것을 보여준다. 확산 모델을 호프필드 네트워크와 유사한 에너지 기반 시스템으로 프레임워크화함으로써, 수렴성, 기억 용량, 최적화에 대한 이론적 보장을 가능하게 하며, 생성 모델링과 기억 회상 기반의 동적 시스템 프레임워크를 통합한다.

ABSTRACT

The generative process of Diffusion Models (DMs) has recently set state-of-the-art on many AI generation benchmarks. Though the generative process is traditionally understood as an "iterative denoiser", there is no universally accepted language to describe it. We introduce a novel perspective to describe DMs using the mathematical language of memory retrieval from the field of energy-based Associative Memories (AMs), making efforts to keep our presentation approachable to newcomers to both of these fields. Unifying these two fields provides insight that DMs can be seen as a particular kind of AM where Lyapunov stability guarantees are bypassed by intelligently engineering the dynamics (i.e., the noise and step size schedules) of the denoising process. Finally, we present a growing body of evidence that records DMs exhibiting empirical behavior we would expect from AMs, and conclude by discussing research opportunities that are revealed by understanding DMs as a form of energy-based memory.

연구 동기 및 목표

  • 동역학 시스템과 상미분방정식을 사용하여 확산 모델과 연상 기억 사이의 공식적인 수학적 연결을 수립하기.
  • 확산 모델의 역방향 노이즈 제거 과정이 연상 기억의 에너지 경사 하강법과 동치임을 보여주기.
  • 표준 확산 모델에서 부족한 고정점으로의 수렴 보장을 포함한 연상 기억의 이론적 이점들을 드러내기.
  • 확산 모델과 트랜스포머 아키텍처를 연상 기억 이론의 관점에서 재해석함으로써 생성 모델링 분야의 새로운 연구 방향을 제안하기.
  • 에너지 기반 시스템의 기억 용량을 통해 대규모 모델의 스케일링 법칙을 이론적으로 이해하는 기초를 제공하기.

제안 방법

  • 연속 시간 상미분방정식으로서의 확산 모델을 정식화하여, 리아푸노프 에너지 함수에 대한 경사 하강법과의 동치성을 드러내기.
  • 확산 모델의 역방향 노이즈 제거 과정을 연상 기억의 기억 회상 역학으로 매핑하여, 에너지 경관의 국소 최소점으로 수렴함을 보여주기.
  • 스코어 기반 모델에서 유도된 에너지 함수를 바탕으로 리아푸노프 함수를 정의하여, 역과정에서의 안정성과 수렴성을 보장하기.
  • 딥러닝에서 널리 사용되는 최적화 기법(예: ADAM, L-BFGS)을 확산 모델의 추론 과정에 적용하기 위해, 이를 에너지에 대한 경사 하강법으로 간주하기.
  • 트랜스포머에 대한 유사성을 확장하여, 어텐션 메커니즘이 조밀한 연상 기억의 단일 스텝 업데이트와 유사함을 보여주고, 이를 바탕으로 '에너지 트랜스포머' 블록을 제안하기.
  • 에너지 경관 내 저장된 고정점의 수를 기반으로 기억 용량과 스케일링 법칙을 분석하여, 모델 크기와 성능을 연상 기억 용량과 연결하기.
Figure 1: Comparing the emphases of Diffusion Models and Associative Memories tasked with learning the same energy (negative log-probability) landscape, represented with both contours and gradient arrows. Diffusion Models (left) train a score function (depicted as orange arrows) to model the gradien
Figure 1: Comparing the emphases of Diffusion Models and Associative Memories tasked with learning the same energy (negative log-probability) landscape, represented with both contours and gradient arrows. Diffusion Models (left) train a score function (depicted as orange arrows) to model the gradien

실험 결과

연구 질문

  • RQ1확산 모델의 역방향 노이즈 제거 역학은 연상 기억의 기억 회상 역학과 어떻게 관련이 있는가?
  • RQ2연상 기억에는 내재되어 있으나 표준 확산 모델에는 부족한 이론적 보장(예: 수렴성)은 무엇인가?
  • RQ3확산 모델의 추론 과정을 에너지 함수에 대한 경사 하강법으로 재해석할 수 있으며, 그로 인해 어떤 최적화 이점이 발생하는가?
  • RQ4최신 아키텍처인 트랜스포머가 그들의 계산 역학에서 얼마나 연상 기억과 유사한가?
  • RQ5대규모 모델의 기억 용량과 스케일링 법칙은 연상 기억 이론의 관점에서 어떻게 이론적으로 이해할 수 있는가?

주요 결과

  • 확산 모델의 역방향 노이즈 제거 과정은 연상 기억과 마찬가지로 리아푸노프 에너지 함수에 대한 경사 하강법과 수학적으로 동치이다.
  • 확산 모델와 달리, 연상 기억은 고정점(국소 최소점)으로의 수렴을 보장하므로, 이론적 안정성 측면에서 우월한 이점이 있다.
  • 연상 기억의 에너지 함수는 직접 계산되고 최적화될 수 있어, 표준 확산 모델에서는 확보할 수 없는 새로운 분석 및 훈련 기법을 가능하게 한다.
  • 확산 모델은 고정점으로 수렴하는 연속 시간 동역학 시스템으로 재해석될 수 있으며, 이로 인해 기억 용량과 안정성의 이론적 특성화가 가능해진다.
  • 확산 모델의 추론 과정은 에너지 경사 하강법으로 간주함으로써, ADAM, L-BFGS와 같은 고급 최적화 기법을 적용해 가속화할 수 있다.
  • 트랜스포머의 어텐션 메커니즘은 조밀한 연상 기억의 단일 스텝 업데이트와 유사하게 작동하며, 어텐션과 기억 회상 간의 더 깊은 이론적 연결을 시사한다.
Figure 2: Diffusion Models through the lens of PF-ODEs, adapted from [ 2 ] . The forward (corrupting) process takes a true data point at time $s=0$ and corrupts the data into pure noise at time $s=T$ using an SDE. The reverse (reconstructing) process undoes the corruption using the score of the dist
Figure 2: Diffusion Models through the lens of PF-ODEs, adapted from [ 2 ] . The forward (corrupting) process takes a true data point at time $s=0$ and corrupts the data into pure noise at time $s=T$ using an SDE. The reverse (reconstructing) process undoes the corruption using the score of the dist

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.