Skip to main content
QUICK REVIEW

[논문 리뷰] Heterogeneity-Aware Asynchronous Decentralized Training

Qinyi Luo, Jiaao He|arXiv (Cornell University)|2019. 09. 17.
Stochastic Gradient Optimization Techniques참고 문헌 50인용 수 8
한 줄 요약

Ripples는 저비용 동기화를 위한 Partial All-Reduce와 충돌을 고려한 그룹 스케줄링 기술(그룹 버퍼, 그룹 분할, 정적/스마트 스케줄링)을 결합한 이질성 인지 비동기 분산 학습 프레임워크이다. 동질적 환경에서는 최신 기술인 All-Reduce 대비 1.1배 빠르며, 이질적 환경에서는 All-Reduce 대비 2배 빠르며, AD-PSGD 대비 이질성 내성에서 3배 빠르다.

ABSTRACT

Distributed deep learning training usually adopts All-Reduce as the synchronization mechanism for data parallel algorithms due to its high performance in homogeneous environment. However, its performance is bounded by the slowest worker among all workers, and is significantly slower in heterogeneous situations. AD-PSGD, a newly proposed synchronization method which provides numerically fast convergence and heterogeneity tolerance, suffers from deadlock issues and high synchronization overhead. Is it possible to get the best of both worlds - designing a distributed training method that has both high performance as All-Reduce in homogeneous environment and good heterogeneity tolerance as AD-PSGD? In this paper, we propose Ripples, a high-performance heterogeneity-aware asynchronous decentralized training approach. We achieve the above goal with intensive synchronization optimization, emphasizing the interplay between algorithm and system implementation. To reduce synchronization cost, we propose a novel communication primitive Partial All-Reduce that allows a large group of workers to synchronize quickly. To reduce synchronization conflict, we propose static group scheduling in homogeneous environment and simple techniques (Group Buffer and Group Division) to avoid conflicts with slightly reduced randomness. Our experiments show that in homogeneous environment, Ripples is 1.1 times faster than the state-of-the-art implementation of All-Reduce, 5.1 times faster than Parameter Server and 4.3 times faster than AD-PSGD. In a heterogeneous setting, Ripples shows 2 times speedup over All-Reduce, and still obtains 3 times speedup over the Parameter Server baseline.

연구 동기 및 목표

  • 가장 느린 워커에 의한 제한으로 인해 이질적 분산 학습 환경에서 All-Reduce의 성능 저하 문제 해결.
  • AD-PSGD에서 발생하는 정지 상태 및 높은 동기화 오버헤드 문제를 해결하면서도 이질성 내성 유지.
  • 동질적 환경에서는 All-Reduce 수준의 성능을, 이질적 환경에서는 AD-PSGD 수준의 내성 확보를 위한 탈중앙화 학습 시스템 설계.
  • 알고리즘 설계와 시스템 수준의 통신 프리미티브 간의 상호작용 최적화를 통해 동기화 비용과 충돌 최소화.

제안 방법

  • 통신 볼륨을 줄여 대규모 워커 그룹 간 빠른 동기화를 가능하게 하는 새로운 통신 프리미티브인 Partial All-Reduce 도입.
  • 무작위성 감소와 동기화 충돌 감소를 위해 동질적 환경에서 정적 그룹 스케줄링 구현.
  • 최소한의 무작위성 손실로 비동기 탈중앙화 학습에서 충돌 확률을 감소시키기 위한 그룹 버퍼 및 그룹 분할 기법 제안.
  • 이질적 환경에서 통계적 효율성과 실행 속도의 균형을 이루기 위해 무작위 그룹 선택(random GG)과 스마트 그룹 스케줄링(smart GG) 사용.
  • 충돌 회피와 함께 원자적 업데이트 프로토콜 통합으로 수치 안정성을 유지하면서 처리량 향상.
  • 유연한 통신 그래프, 스펙트럼 갭, 이중 스토하스틱 평균화 성질을 갖춘 탈중앙화 학습 시스템 설계로 강건하고 확장 가능한 학습 지원.

실험 결과

연구 질문

  • RQ1탈중앙화 학습 시스템이 동질적 환경에서는 All-Reduce 수준의 성능를 달성하면서도 AD-PSGD 수준의 이질성 내성 유지가 가능한가?
  • RQ2수렴 안정성을 훼손하지 않고 대규모 탈중앙화 학습에서 동기화 비용을 어떻게 최소화할 수 있는가?
  • RQ3그룹 선택의 무작위성과 동기화 충돌 간의 상충 관계는 무엇이며, 어떻게 최적화할 수 있는가?
  • RQ4충돌 방지 기법이 이질적 워커 환경에서 성능 향상에 얼마나 기여하는가?

주요 결과

  • 동질적 환경에서 Ripples는 최신 All-Reduce 구현 대비 1.1배 빠르며, 파라미터 서버 대비 5.1배, AD-PSGD 대비 4.3배 빠르다.
  • 한 워커에서 2배 속도 저하가 발생하는 이질적 환경에서 Ripples는 All-Reduce 대비 2배 빠르며, 파라미터 서버 대비 3배 빠르다.
  • 5배 속도 저하 상황에서도 Ripples는 All-Reduce 대비 5.01배 빠르고, AD-PSGD 대비 4.23배 빠르며, 강력한 이질성 내성 입증.
  • Ripples의 스마트 그룹 스케줄링은 AD-PSGD 대비 동기화 오버헤드 감소와 더 나은 충돌 관리 덕분에 총 실행 시간에서 5.26배 빠르다.
  • ImageNet에서 ResNet-50을 사용할 경우, Ripples는 스마트 그룹 스케줄링로 10시간 후에 64.21%의 Top-1 정확도를 달성하여 AD-PSGD(58.28%)를 능가하고, All-Reduce(66.83%)에 수렴하는 데서도 반복 횟수가 더 많음에도 불구하고 수렴 속도에서 앞서며 성능 우월성 입증.
  • Ripples의 성능 향상은 반복당 동기화 비용을 5.10배 빠르게 한 데 기인하며, 이는 약간 더 많은 반복 횟수(0.96× 대비 0.78×)를 감수함으로써 효율성과 통계적 수렴 간 효과적인 트레이드오프를 입증.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.