Skip to main content
QUICK REVIEW

[논문 리뷰] Theoretical Foundations of Deep Selective State-Space Models

Nicola Muca Cirone, Antonio Orvieto|arXiv (Cornell University)|2024. 02. 29.
Complex Systems and Decision Making인용 수 4
한 줄 요약

이 논문은 리프 경로 이론을 사용하여 딥 선택적 상태공간모델(SSI)의 이론적 기반을 수립하며, Mamba와 GLA와 같은 모델에서 입력 제어 선형 재귀관계가 입력의 서명에 대한 투영을 통해 시간적 스케일 간 비선형 상호작용을 확실히 포착할 수 있음을 증명한다. 핵심 결과는 넓은 크기의 무작위로 초기화된 선택적 SSM이 추가적인 MLP가 없이도 완전 표현 가능하다는 것이며, 대각선 형태의 변형은 깊이를 통해 고차원 통계를 효율적으로 계산한다.

ABSTRACT

Structured state-space models (SSMs) such as S4, stemming from the seminal work of Gu et al., are gaining popularity as effective approaches for modeling sequential data. Deep SSMs demonstrate outstanding performance across a diverse set of domains, at a reduced training and inference cost compared to attention-based transformers. Recent developments show that if the linear recurrence powering SSMs allows for multiplicative interactions between inputs and hidden states (e.g. GateLoop, Mamba, GLA), then the resulting architecture can surpass in both in accuracy and efficiency attention-powered foundation models trained on text, at scales of billion parameters. In this paper, we give theoretical grounding to this recent finding using tools from Rough Path Theory: we show that when random linear recurrences are equipped with simple input-controlled transitions (selectivity mechanism), then the hidden state is provably a low-dimensional projection of a powerful mathematical object called the signature of the input -- capturing non-linear interactions between tokens at distinct timescales. Our theory not only motivates the success of modern selective state-space models such as Mamba but also provides a solid framework to understand the expressive power of future SSM variants.

연구 동기 및 목표

  • 선택적 SSM의 현대적 분석을 위해 리프 경로 이론의 도구를 활용한 이론적 프레임워크를 제공하는 것.
  • SSM 내 입력 제어 선형 재귀관계가 입력 시퀀스의 연속 함수를 근사하는 데 있어 완전 표현 가능하다는 것을 증명하는 것.
  • 대각선 SSM(예: Mamba)이 깊이를 통해 고차원 통계를 효율적으로 계산함으로써 밀도 있는 변형과 동일한 표현력을 가지는지 보여주는 것.
  • 제어 미분 방정식(CDE) 기반의 공통 수학적 공식을 통해 S4, Mamba, GLA 및 관련 모델의 분석을 통합하는 것.
  • 향후 SSM의 다양한 변형을 더 높은 표현력과 효율성으로 이해하고 설계할 수 있도록 일반화 가능한 이론적 기초를 제공하는 것.

제안 방법

  • 입력 시퀀스에 의해 구동되는 제어 미분 방정식(CDE)으로 SSM을 모델링하며, 은닉 상태는 입력에 의존하는 선형 전이를 통해 진화한다.
  • 리프 경로 이론을 적용하여 은닉 상태를 입력 서명의 저차원 투영으로 형식화함으로써 비선형 시간적 상호작용을 포착한다.
  • 상태 전이 행렬 $ W_{s,t} $ 를 입력 유도 벡터장의 반복 적분의 합으로 정의하여 S4 및 Mamba 재귀관계를 일반화한다.
  • 와른스키안 행렬과 경로 적분 표현을 사용하여 덧셈형 입력 항이 있는 CDE의 명시적 해를 유도한다.
  • 무작위로 초기화된 밀도 있는 입력 제어 재귀관계가 은닉 상태에서 충분한 입력 통계를 포착함으로써 완전 표현 가능하다는 것을 증명한다.
  • 대각선 SSM 블록을 연결함으로써 고차원 전역 통계를 계산하며, 깊이의 극한에서 밀도 있는 모델과 동등한 표현력을 달성함을 보여준다.
Figure 1 : Example of the dynamics of the S4 and Mamba CDEs derived in Sec. 3.1 . The continuous-time approximation of these algorithms can be written as a Linear CDE. In S4, $\omega^{\text{X}}_{t}=t$ , while $\xi^{\text{X}}_{t}=\int_{0}^{t}X_{s}ds$ . In Mamba, $\omega^{\text{X}}_{t}=\int_{0}^{t}\si
Figure 1 : Example of the dynamics of the S4 and Mamba CDEs derived in Sec. 3.1 . The continuous-time approximation of these algorithms can be written as a Linear CDE. In S4, $\omega^{\text{X}}_{t}=t$ , while $\xi^{\text{X}}_{t}=\int_{0}^{t}X_{s}ds$ . In Mamba, $\omega^{\text{X}}_{t}=\int_{0}^{t}\si

실험 결과

연구 질문

  • RQ1Mamba와 같은 현대 선택적 SSM의 표현력은 어떻게 이론적으로 기반을 두어야 하는가?
  • RQ2선택적 SSM이 장거리 시퀀스에서 비선형 상호작용을 모델링할 수 있는 데 기여하는 수학적 구조는 무엇인가?
  • RQ3추가적인 비선형 구성 요소(예: MLP) 없이도 입력 제어 선형 재귀관계가 완전 표현 가능할 수 있는가?
  • RQ4대각선 SSM은 고차원 입력 통계를 포착하는 데서 밀도 있는 변형과 어떻게 비교되는가?
  • RQ5제어 미분 방정식과 리프 경로 이론을 기반으로 한 통합 프레임워크를 통해 다양한 SSM 아키텍처를 분석하고 비교할 수 있는가?

주요 결과

  • 선택적 SSM 내 입력 제어 선형 재귀관계는 입력 서명의 저차원 투영으로 은닉 상태를 확실히 표현하며, 비선형 시간적 상호작용을 포착한다.
  • 무작위로 초기화된 넓은 크기의 밀도 있는 입력 제어 SSM은 완전 표현 가능하다: 추가적인 MLP가 없이도 입력 시퀀스에서 출력으로의 임의의 연속 함수를 근사할 수 있다.
  • Mamba와 같은 대각선 SSM은 깊이를 통해 고차원 통계를 효율적으로 계산하며, 깊이의 극한에서 밀도 있는 모델과 동등한 표현력을 달성한다.
  • 선택적 SSM의 은닉 상태는 수학적으로 제어 미분 방정식(CDE)을 푸는 것과 동일하며, 경로 적분 표현과 와른스키안 행렬을 통해 명시적 해가 도출된다.
  • 이 프레임워크는 S4, Mamba, GLA 및 관련 모델을 동일한 이론적 기초 아래 통합하여 향후 SSM 변형의 엄밀한 비교 및 분석을 가능하게 한다.
  • 이 이론은 선택적 SSM이 표준 S4와 어텐션 메커니즘보다 장거리 모델링에서 뛰어나게 작용하는 이유를 설명한다: 구조적이고 입력에 의존하는 전이를 통해 더 풍부한 입력 통계를 포착하기 때문이다.
Figure 3 : Comparison of Linear CDE, Mamba and S5 on anti-symmetric signature prediction tasks. As suggested by the theory Mamba struggles to immediately capture these high order statistics, but performance improves with chaining . On the other hand S5 does not learn such non-linear features even af
Figure 3 : Comparison of Linear CDE, Mamba and S5 on anti-symmetric signature prediction tasks. As suggested by the theory Mamba struggles to immediately capture these high order statistics, but performance improves with chaining . On the other hand S5 does not learn such non-linear features even af

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.