Skip to main content
QUICK REVIEW

[논문 리뷰] ORIGAMI: A Heterogeneous Split Architecture for In-Memory Acceleration of Learning

Hajar Falahati, Pejman Lotfi-Kamran|arXiv (Cornell University)|2018. 12. 30.
Parallel Computing and Optimization Techniques참고 문헌 69인용 수 5
한 줄 요약

Origami는 기계학습 워크로드를 위한 3D스택드 메모리의 대역폭, 전력, 면적 제약를 극복하기 위해 메모리 내 가속기와 외부 계산 플랫폼을 조합한 이종 분할 아키텍처를 제안한다. 공통된 계산 패턴을 전용 엔진으로 추출하고, 메모리 내 및 외부 플랫폼 간에 지능적으로 계산을 분할함으로써, 기존 최고 수준의 가속기 대비 최대 1.6배의 성능 향상과 31배 향상된 에너지-지연 제품(Energy-Delay Product, EDP)을 달성한다.

ABSTRACT

Memory bandwidth bottleneck is a major challenges in processing machine learning (ML) algorithms. In-memory acceleration has potential to address this problem; however, it needs to address two challenges. First, in-memory accelerator should be general enough to support a large set of different ML algorithms. Second, it should be efficient enough to utilize bandwidth while meeting limited power and area budgets of logic layer of a 3D-stacked memory. We observe that previous work fails to simultaneously address both challenges. We propose ORIGAMI, a heterogeneous set of in-memory accelerators, to support compute demands of different ML algorithms, and also uses an off-the-shelf compute platform (e.g.,FPGA,GPU,TPU,etc.) to utilize bandwidth without violating strict area and power budgets. ORIGAMI offers a pattern-matching technique to identify similar computation patterns of ML algorithms and extracts a compute engine for each pattern. These compute engines constitute heterogeneous accelerators integrated on logic layer of a 3D-stacked memory. Combination of these compute engines can execute any type of ML algorithms. To utilize available bandwidth without violating area and power budgets of logic layer, ORIGAMI comes with a computation-splitting compiler that divides an ML algorithm between in-memory accelerators and an out-of-the-memory platform in a balanced way and with minimum inter-communications. Combination of pattern matching and split execution offers a new design point for acceleration of ML algorithms. Evaluation results across 12 popular ML algorithms show that ORIGAMI outperforms state-of-the-art accelerator with 3D-stacked memory in terms of performance and energy-delay product (EDP) by 1.5x and 29x (up to 1.6x and 31x), respectively. Furthermore, results are within a 1% margin of an ideal system that has unlimited compute resources on logic layer of a 3D-stacked memory.

연구 동기 및 목표

  • 3D스택드 메모리를 활용하여 기계학습 훈련에서의 메모리 대역폭 병목을 해결한다.
  • 일반성 부족 또는 가용 대역폭를 효율적으로 활용하지 못하는 기존 메모리 내 가속기의 한계를 극복한다.
  • 3D스택드 메모리의 논리 레이어에 있는 엄격한 전력 및 면적 예산을 고려하면서 다양한 기계학습 알고리즘을 지원하는 시스템을 설계한다.
  • 지능적인 계산 분할을 통해 메모리 내 가속기와 외부 계산 플랫폼을 조합하여 메모리 대역폭을 전면적으로 활용한다.

제안 방법

  • 패턴 매칭을 사용하여 12개의 기계학습 알고리즘 간의 공통된 계산 패턴을 식별하고, 전용이면서 저비용의 계산 엔진을 유도한다.
  • 이러한 이종 계산 엔진을 3D스택드 메모리의 논리 레이어에 통합하여 다양한 기계학습 알고리즘을 지원한다.
  • 메모리 내 가속기와 외부 플랫폼(FPGA, GPU, TPU 등) 간에 기계학습 워크로드를 분할하여 부하 균형을 맞추고 플랫폼 간 통신을 최소화하는 계산 분할 컴파일러를 개발한다.
  • 칩 내 가속기(최대 47% 대역폭 활용)와 칩 외부 플랫폼 액세스(내부 대역폭의 최대 63%)를 조합하여 3D스택드 메모리의 내부 대역폭을 효율적으로 활용한다.
  • 블록 수준, 부분 수준, 모델 수준의 3단계 병렬 처리를 적용하여 워크로드 분포 및 자원 활용도를 동적으로 최적화한다.

실험 결과

연구 질문

  • RQ1엄격한 전력 및 면적 제약 내에서 이종 메모리 내 가속기 조합이 다양한 기계학습 알고리즘을 효율적으로 지원할 수 있는가?
  • RQ2메모리 내 가속기만으로 3D스택드 DRAM 아키텍처에서 가용 메모리 대역폭을 얼마나 효율적으로 활용할 수 있는가?
  • RQ3메모리 내 가속기와 외부 계산 플랫폼 간에 계산을 효과적으로 분할하여 대역폭 활용도를 극대화하고 통신 오버헤드를 최소화할 수 있는가?
  • RQ4패턴 기반 메모리 내 가속과 외부 실행을 조합함으로써 달성 가능한 성능 및 에너지 효율성 향상 수준은 어느 정도인가?

주요 결과

  • Origami는 12개의 기계학습 벤치마크에서 기존 최고 수준의 3D스택드 메모리 가속기 대비 최대 1.6배 성능 향상과 31배 향상된 에너지-지연 제품(EDP)을 달성한다.
  • FPGA 기반 가속 대비 평균 1.5배의 성능 향상과 21배 향상된 EDP를 기록하여 뛰어난 확장성과 효율성을 입증한다.
  • 모델 수준 병렬 처리가 가능한 벤치마크(예: Reco)의 경우, 블록 수준, 부분 수준, 모델 수준의 병렬 처리를 모두 활성화하면 최대 1.6배 성능 향상과 31배 EDP 감소를 달성한다.
  • 무제한 계산 자원이 있는 이상적인 시스템과 비교해 1% 이내의 성능를 기록하여 거의 최적의 대역폭 활용도를 입증한다.
  • 부분 수준 및 모델 수준 병렬 처리를 활성화하면 복잡한 모델(예: BProp, Reco)의 경우 성능과 EDP에 상당한 향상이 있으며, 블록 수준 병렬 처리만 사용할 경우 대비 최대 26.75%의 EDP 감소를 기록한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.