Skip to main content
QUICK REVIEW

[논문 리뷰] STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning

Prashant Khanduri, Pranay Sharma|arXiv (Cornell University)|2021. 06. 19.
Privacy-Preserving Technologies in Data참고 문헌 35인용 수 20
한 줄 요약

이 논문은 비볼록 설정에서 $\epsilon$-정류해를 찾는 데 near-optimal 샘플 복잡도 $\tilde{\mathcal{O}}(\epsilon^{-3/2})$와 통신 복잡도 $\tilde{\mathcal{O}}(\epsilon^{-1})$를 달성하는 분산 학습을 위한 확률적 이면형 모멘타움 알고리즘인 STEM을 제안한다. 적응형 미니배치 크기와 국소 업데이트 빈도를 함께 최적화함으로써 워커와 서버의 모멘타움 기반 업데이트를 조정함으로써, STEM은 이러한 복잡도를 유지하는 트레이드오프 곡선을 수립하며, FedAvg 및 이전 방법들을 능가한다.

ABSTRACT

Federated Learning (FL) refers to the paradigm where multiple worker nodes (WNs) build a joint model by using local data. Despite extensive research, for a generic non-convex FL problem, it is not clear, how to choose the WNs' and the server's update directions, the minibatch sizes, and the local update frequency, so that the WNs use the minimum number of samples and communication rounds to achieve the desired solution. This work addresses the above question and considers a class of stochastic algorithms where the WNs perform a few local updates before communication. We show that when both the WN's and the server's directions are chosen based on a stochastic momentum estimator, the algorithm requires $ ilde{\mathcal{O}}(ε^{-3/2})$ samples and $ ilde{\mathcal{O}}(ε^{-1})$ communication rounds to compute an $ε$-stationary solution. To the best of our knowledge, this is the first FL algorithm that achieves such {\it near-optimal} sample and communication complexities simultaneously. Further, we show that there is a trade-off curve between local update frequencies and local minibatch sizes, on which the above sample and communication complexities can be maintained. Finally, we show that for the classical FedAvg (a.k.a. Local SGD, which is a momentum-less special case of the STEM), a similar trade-off curve exists, albeit with worse sample and communication complexities. Our insights on this trade-off provides guidelines for choosing the four important design elements for FL algorithms, the update frequency, directions, and minibatch sizes to achieve the best performance.

연구 동기 및 목표

  • 비볼록 분산 학습에서 워커와 서버의 업데이트 방향, 미니배치 크기, 국소 업데이트 빈도를 동시에 최적화하는 열린 문제를 해결하기 위해.
  • 이질적인 데이터를 가진 확률적이고 분산된 환경에서 샘플 복잡도와 통신 복잡도를 동시에 최소화하기 위해.
  • 국소 업데이트 빈도와 미니배치 크기 사이의 이론적 트레이드오프를 수립하여 near-optimal 수렴 속도를 유지하기 위해.
  • FedAvg(Local SGD)에 대한 새로운 통찰을 제공하여 그 내재된 트레이드오프와 비최적의 복잡도 한계를 밝혀내기 위해.
  • 기본 가정 하에 워커와 서버 양쪽에서 모멘타움 기반 업데이트가 최적의 수렴을 가능하게 함을 보여주기 위해.

제안 방법

  • 워커 노드와 서버가 모두 모멘타움 추정 기반 그라디언트를 사용하여 모델 업데이트를 수행하는 확률적 이면형 모멘타움 알고리즘인 STEM을 제안한다.
  • 국소 업데이트 빈도 $I$, 미니배치 크기 $b$, 모멘타움 기반 업데이트 방향을 함께 최적화하는 통합 프레임워크를 도입한다.
  • 모멘타움을 통한 분산 감소를 활용한 확률적 그라디언트 추정기를 활용하여 학습을 안정화하고 그라디언트 노이즈를 감소시킨다.
  • 비볼록 스무쓰함과 유한한 그라디언트 가정 하에 리아파노프 분석과 강凸성 유사 부등식을 사용하여 수렴 한계를 유도한다.
  • STEM이 $\tilde{\mathcal{O}}(\epsilon^{-3/2})$ 샘플과 $\tilde{\mathcal{O}}(\epsilon^{-1})$ 통신 라운드를 달성하도록 하는 $I$와 $b$ 사이의 트레이드오프 곡선을 수립한다.
  • FedAvg를 STEM의 특수 케이스(모멘타움 없음)로 분석하고 비교를 위해 그의 비최적 복잡도 한계를 도출한다.

실험 결과

연구 질문

  • RQ1비볼록 설정에서 분산 학습 알고리즘이 near-optimal 샘플 복잡도와 통신 복잡도를 동시에 달성할 수 있는가?
  • RQ2확률적 그라디언트를 가진 분산 학습에서 국소 업데이트 빈도와 미니배치 크기 사이의 최적 트레이드오프는 무엇인가?
  • RQ3워커와 서버 양쪽에 모멘타움을 통합할 경우 비볼록 FL에서 수렴 복잡도에 어떤 영향을 미치는가?
  • RQ4STEM과 동일한 가정 하에 FedAvg (Local SGD)의 이론적 샘플 및 통신 복잡도는 무엇인가?
  • RQ5최소 자원 사용을 위해 업데이트 방향, 미니배치 크기, 국소 업데이트 빈도를 동시에 최적화할 수 있는 통합 프레임워크를 개발할 수 있는가?

주요 결과

  • STEM은 비볼록 분산 학습에서 $\epsilon$-정류해를 찾는 데 $\tilde{\mathcal{O}}(\epsilon^{-3/2})$ 샘플 복잡도와 $\tilde{\mathcal{O}}(\epsilon^{-1})$ 통신 복잡도를 달성한다.
  • 미니배치 크기 $b$와 국소 업데이트 빈도 $I$ 사이에 트레이드오프 곡선이 존재하며, 이 곡선을 따라 $b$와 $I$를 선택할 경우 STEM은 near-optimal 복잡도를 유지한다.
  • FedAvg (STEM의 모멘타움 없음 버전)의 경우 샘플 복잡도는 $\mathcal{O}(\epsilon^{-2})$, 통신 복잡도는 $\mathcal{O}(\epsilon^{-3/2})$이며, STEM에 비해 비최적이다.
  • 만약 $b = \mathcal{O}(1)$이고 $I = (T/b^3K^3)^{1/4}$일 경우, STEM은 $\tilde{\mathcal{O}}(\epsilon^{-3/2})$ 샘플 복잡도와 $\tilde{\mathcal{O}}(\epsilon^{-1})$ 통신 복잡도를 유지한다.
  • 분석 결과, 워커와 서버 양쪽에서 모멘타움이 확률적 분산 학습에서 최적의 수렴 속도를 달성하는 데 필수적임을 밝혀냈다.
  • 유도된 한계는 표준 가정 하에 날카롭게 조밀하며, STEM은 이 설정에서 동시에 near-optimal 샘플 및 통신 복잡도를 달성하는 최초의 알고리즘이다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.