Skip to main content
QUICK REVIEW

[논문 리뷰] Jamba: A Hybrid Transformer-Mamba Language Model

Opher Lieber, Barak Lenz|arXiv (Cornell University)|2024. 03. 28.
Language, Discourse, Communication Strategies인용 수 40
한 줄 요약

Jamba는 Transformer와 Mamba 계층을 MoE와 교대로 배치하는 하이브리드 Transformer-Mamba 혼합 전문가 아키텍처를 도입하여 단일 80GB GPU에 적합하면서 높은 성능과 긴 컨텍스트 능력을 달성합니다.

ABSTRACT

We present Jamba, a new base large language model based on a novel hybrid Transformer-Mamba mixture-of-experts (MoE) architecture. Specifically, Jamba interleaves blocks of Transformer and Mamba layers, enjoying the benefits of both model families. MoE is added in some of these layers to increase model capacity while keeping active parameter usage manageable. This flexible architecture allows resource- and objective-specific configurations. In the particular configuration we have implemented, we end up with a powerful model that fits in a single 80GB GPU. Built at large scale, Jamba provides high throughput and small memory footprint compared to vanilla Transformers, and at the same time state-of-the-art performance on standard language model benchmarks and long-context evaluations. Remarkably, the model presents strong results for up to 256K tokens context length. We study various architectural decisions, such as how to combine Transformer and Mamba layers, and how to mix experts, and show that some of them are crucial in large scale modeling. We also describe several interesting properties of these architectures which the training and evaluation of Jamba have revealed, and plan to release checkpoints from various ablation runs, to encourage further exploration of this novel architecture. We make the weights of our implementation of Jamba publicly available under a permissive license.

연구 동기 및 목표

  • Transformer와 Mamba 계층을 교대로 배치하는 것이 주의 기반 모델과 상태 공간 모델의 강점을 결합할 수 있는지 조사한다.
  • 하이브리드 아키텍처에서 MoE 통합이 용량, 처리량 및 메모리에 어떤 영향을 미치는지 평가한다.
  • 256K 토큰 컨텍스트를 포함한 표준 벤치마크 및 긴 컨텍스트 작업에서의 성능을 평가한다.
  • 7B-12B 매개변수 규모의 하이브드 모델을 커머더티 하드웨어에서 학습 안정성과 실용성을 시연한다.

제안 방법

  • Jamba 블록은 Transformer 또는 Mamba 계층 뒤에 MLP 또는 MoE 모듈을 결합하도록 정의한다.
  • a:m 주의-대 Mamba 비율로 블록을 교대로 구성하고, e 계층마다 MoE를 적용하며 총 n개의 전문가와 토큰당 top-K 라우팅을 적용한다.
  • Mamba 계층에서 RMSNorm을 사용하고 명시적 위치 임베딩을 생략하며 하이브리드 구조가 암시적 위치 정보를 제공하도록 한다.
  • 64K 어휘와 BPE 토크나이저로 대규모 데이터에 대해 학습하며 80GB GPU 설정에서 처리량 및 메모리 효율성을 최적화한다.
  • 학술 벤치마크, 긴 컨텍스트 QA 데이터셋, 다양한 하드웨어 규모에서 처리량 측정을 평가한다.

실험 결과

연구 질문

  • RQ1동일한 규모의 순수 Transformer 모델과 비교하여 하이브리드 Attention-Mamba 아키텍처가 표준 벤치마크에서 이를 능가하거나 따라잡을 수 있는가?
  • RQ2하이브리드 아키텍처에 MoE를 도입하면 무거운 계산 없이 용량이 향상되는가?
  • RQ3Attention 계층과 Mamba 계층의 비율이 메모리 사용량, 처리량, 긴 컨텍스트 성능에 어떤 영향을 미치는가?
  • RQ4Jamba가 매우 긴 컨텍스트(최대 256K 토큰)를 합리적인 KV 캐시 요구사항으로 효과적으로 처리할 수 있는가?
  • RQ5대규모 하이브리드 모델의 실용적 학습 안정성 고려사항은 무엇인가?

주요 결과

  • Jamba는 Mixtral 및 Llama-2 70B와 같은 유사 규모의 공개 모델과 비교하여 표준 벤치마크에서 경쟁력 있거나 우수한 정확도를 달성한다.
  • 하이브리드 Attention-Mamba 아키텍처는 256K 컨텍스트에서 KV 캐시 요구사항을 4GB로 낮추어 단일 80GB GPU에서 긴 컨텍스트 처리를 가능하게 한다.
  • MoE 변형은 대규모 규모에서 비-MoE 하이브리드보다 성능을 향상시키며(50B 토큰에 대해 7B 매개변수 학습 시) 실험이 뒷받침된다.
  • Attention-Mamba 하이브리드는 여러 작업에서 순수 Mamba보다 우수하며 Transformers와 유사한 인-컨텍스트 학습을 지원하여 상호 보완적 강점을 시사한다.
  • Jamba는 명시적 위치 정보를 필요로 하지 않으며, Mamba-우선 구조가 암시적 위치 정보를 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.