[논문 리뷰] ORIGAMI: A Heterogeneous Split Architecture for In-Memory Acceleration of Learning
Origami는 기계학습 워크로드를 위한 3D스택드 메모리의 대역폭, 전력, 면적 제약를 극복하기 위해 메모리 내 가속기와 외부 계산 플랫폼을 조합한 이종 분할 아키텍처를 제안한다. 공통된 계산 패턴을 전용 엔진으로 추출하고, 메모리 내 및 외부 플랫폼 간에 지능적으로 계산을 분할함으로써, 기존 최고 수준의 가속기 대비 최대 1.6배의 성능 향상과 31배 향상된 에너지-지연 제품(Energy-Delay Product, EDP)을 달성한다.
Memory bandwidth bottleneck is a major challenges in processing machine learning (ML) algorithms. In-memory acceleration has potential to address this problem; however, it needs to address two challenges. First, in-memory accelerator should be general enough to support a large set of different ML algorithms. Second, it should be efficient enough to utilize bandwidth while meeting limited power and area budgets of logic layer of a 3D-stacked memory. We observe that previous work fails to simultaneously address both challenges. We propose ORIGAMI, a heterogeneous set of in-memory accelerators, to support compute demands of different ML algorithms, and also uses an off-the-shelf compute platform (e.g.,FPGA,GPU,TPU,etc.) to utilize bandwidth without violating strict area and power budgets. ORIGAMI offers a pattern-matching technique to identify similar computation patterns of ML algorithms and extracts a compute engine for each pattern. These compute engines constitute heterogeneous accelerators integrated on logic layer of a 3D-stacked memory. Combination of these compute engines can execute any type of ML algorithms. To utilize available bandwidth without violating area and power budgets of logic layer, ORIGAMI comes with a computation-splitting compiler that divides an ML algorithm between in-memory accelerators and an out-of-the-memory platform in a balanced way and with minimum inter-communications. Combination of pattern matching and split execution offers a new design point for acceleration of ML algorithms. Evaluation results across 12 popular ML algorithms show that ORIGAMI outperforms state-of-the-art accelerator with 3D-stacked memory in terms of performance and energy-delay product (EDP) by 1.5x and 29x (up to 1.6x and 31x), respectively. Furthermore, results are within a 1% margin of an ideal system that has unlimited compute resources on logic layer of a 3D-stacked memory.
연구 동기 및 목표
- 3D스택드 메모리를 활용하여 기계학습 훈련에서의 메모리 대역폭 병목을 해결한다.
- 일반성 부족 또는 가용 대역폭를 효율적으로 활용하지 못하는 기존 메모리 내 가속기의 한계를 극복한다.
- 3D스택드 메모리의 논리 레이어에 있는 엄격한 전력 및 면적 예산을 고려하면서 다양한 기계학습 알고리즘을 지원하는 시스템을 설계한다.
- 지능적인 계산 분할을 통해 메모리 내 가속기와 외부 계산 플랫폼을 조합하여 메모리 대역폭을 전면적으로 활용한다.
제안 방법
- 패턴 매칭을 사용하여 12개의 기계학습 알고리즘 간의 공통된 계산 패턴을 식별하고, 전용이면서 저비용의 계산 엔진을 유도한다.
- 이러한 이종 계산 엔진을 3D스택드 메모리의 논리 레이어에 통합하여 다양한 기계학습 알고리즘을 지원한다.
- 메모리 내 가속기와 외부 플랫폼(FPGA, GPU, TPU 등) 간에 기계학습 워크로드를 분할하여 부하 균형을 맞추고 플랫폼 간 통신을 최소화하는 계산 분할 컴파일러를 개발한다.
- 칩 내 가속기(최대 47% 대역폭 활용)와 칩 외부 플랫폼 액세스(내부 대역폭의 최대 63%)를 조합하여 3D스택드 메모리의 내부 대역폭을 효율적으로 활용한다.
- 블록 수준, 부분 수준, 모델 수준의 3단계 병렬 처리를 적용하여 워크로드 분포 및 자원 활용도를 동적으로 최적화한다.
실험 결과
연구 질문
- RQ1엄격한 전력 및 면적 제약 내에서 이종 메모리 내 가속기 조합이 다양한 기계학습 알고리즘을 효율적으로 지원할 수 있는가?
- RQ2메모리 내 가속기만으로 3D스택드 DRAM 아키텍처에서 가용 메모리 대역폭을 얼마나 효율적으로 활용할 수 있는가?
- RQ3메모리 내 가속기와 외부 계산 플랫폼 간에 계산을 효과적으로 분할하여 대역폭 활용도를 극대화하고 통신 오버헤드를 최소화할 수 있는가?
- RQ4패턴 기반 메모리 내 가속과 외부 실행을 조합함으로써 달성 가능한 성능 및 에너지 효율성 향상 수준은 어느 정도인가?
주요 결과
- Origami는 12개의 기계학습 벤치마크에서 기존 최고 수준의 3D스택드 메모리 가속기 대비 최대 1.6배 성능 향상과 31배 향상된 에너지-지연 제품(EDP)을 달성한다.
- FPGA 기반 가속 대비 평균 1.5배의 성능 향상과 21배 향상된 EDP를 기록하여 뛰어난 확장성과 효율성을 입증한다.
- 모델 수준 병렬 처리가 가능한 벤치마크(예: Reco)의 경우, 블록 수준, 부분 수준, 모델 수준의 병렬 처리를 모두 활성화하면 최대 1.6배 성능 향상과 31배 EDP 감소를 달성한다.
- 무제한 계산 자원이 있는 이상적인 시스템과 비교해 1% 이내의 성능를 기록하여 거의 최적의 대역폭 활용도를 입증한다.
- 부분 수준 및 모델 수준 병렬 처리를 활성화하면 복잡한 모델(예: BProp, Reco)의 경우 성능과 EDP에 상당한 향상이 있으며, 블록 수준 병렬 처리만 사용할 경우 대비 최대 26.75%의 EDP 감소를 기록한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.