Skip to main content
QUICK REVIEW

[논문 리뷰] Acceleration of the Implicit-Explicit Non-hydrostatic Unified Model of the Atmosphere (NUMA) on Manycore Processors

Daniel Abdi, Francis X. Giraldo|arXiv (Cornell University)|2017. 02. 13.
Matrix Theory and Algorithms참고 문헌 28인용 수 6
한 줄 요약

이 논문은 GPU 및 인텔 KNL과 같은 멀티코어 프로세서에서 고급 수치 기법을 활용하여 IMEX 비압축 대기 모델(NUMA)의 성능을 가속화한다. 이 기법들은 슈어 보조식 감소, HEVI 1D-IMEX 설정, 고차수 조절자, 직접 해법기를 포함한다. 고 쿠르랑트 수치에서 명시적 방법 대비 최대 100배의 성능 향상을 달성하며, 엑사스케일 대응 아키텍처에서 뛰어난 확장성과 성능을 입증한다.

ABSTRACT

We present the acceleration of an IMplicit-EXplicit (IMEX) non-hydrostatic atmospheric model on manycore processors such as GPUs and Intel's MIC architecture. IMEX time integration methods sidestep the constraint imposed by the Courant-Friedrichs-Lewy condition on explicit methods through corrective implicit solves within each time step. In this work, we implement and evaluate the performance of IMEX on manycore processors relative to explicit methods. Using 3D-IMEX at Courant number C=15 , we obtained a speedup of about 4X relative to an explicit time stepping method run with the maximum allowable C=1. In addition, we demonstrate a much larger speedup of 100X at C=150 using 1D-IMEX due to the unconditional stability of the method in the vertical direction. Several improvements on the IMEX procedure were necessary in order to outperform our results with explicit methods: a) reducing the number of degrees of freedom of the IMEX formulation by forming the Schur complement; b) formulating a horizontally-explicit vertically-implicit (HEVI) 1D-IMEX scheme that has a lower workload and potentially better scalability than 3D-IMEX; c) using high-order polynomial preconditioners to reduce the condition number of the resulting system; d) using a direct solver for the 1D-IMEX method by performing and storing LU factorizations once to obtain a constant cost for any Courant number. Without all of these improvements, explicit time integration methods turned out to be difficult to beat. We discuss in detail the IMEX infrastructure required for formulating and implementing efficient methods on manycore processors. Finally, we validate our results with standard benchmark problems in NWP and evaluate the performance and scalability of the IMEX method using up to 4192 GPUs and 16 Knights Landing processors.

연구 동기 및 목표

  • 비압축 대기 모델에서 빠른 음파 및 중력파로 인한 명시적 방법의 심각한 시간단계 제약 해결.
  • 다중코어 아키텍처에서 암시적 방법의 성능 저하 문제를 알고리즘 및 구현 최적화로 극복.
  • 해결 정확도를 유지하면서도 큰 쿠르랑트 수치를 허용함으로써 명시적 시간 적분 대비 빠른 성능 향상 달성.
  • 운영용 수치 기상 예측을 위한 현대적 멀티코어 시스템(예: GPU 및 인텔 KNL)에서 고도의 확장성과 성능 확보.
  • 표준 NWP 벤치마크를 사용해 IMEX 접근법을 검증하고, 다양한 하드웨어 플랫폼에서의 강인성 입증.

제안 방법

  • 빠른 파동(음파/중력파)은 암시적으로, 느린 파동(Rossby)은 명시적으로 처리하는 ARK 스킴을 사용한 IMEX 시간 적분 구현.
  • 3D-IMEX의 자유도를 감소시키기 위해 슈어 보조식 설정을 통한 자유도 감소로 계산 효율성 향상.
  • 수평적으로는 명시적, 수직적으로는 암시적인 HEVI 1D-IMEX 스킴 개발으로 작업 부담 감소 및 확장성 향상.
  • 암시적 시스템의 조건수를 감소시키기 위해 고차수 다항식 조절자 적용으로 반복 해법 가속.
  • 1D-IMEX에 대해 사전 계산된 LU 분해를 사용한 직접 해법기를 적용하여 쿠르랑트 수치와 무관한 일정한 비용의 암시적 해법 달성.
  • 모든 커널을 통합된 OCCA 언어로 구현하여 GPU 및 KNL 프로세서에서 이식 가능하고 고성능 실행 구현.

실험 결과

연구 질문

  • RQ1다중코어 프로세서에서 IMEX 방법이 비압축 대기 모델링에서 명시적 방법 대비 뚜렷한 성능 향상을 달성할 수 있는가?
  • RQ2슈어 보조식 및 HEVI 설정과 같은 알고리즘 개선이 GPU 및 KNL에서 IMEX의 성능 및 확장성에 어떤 영향을 미치는가?
  • RQ3쿠르랑트 수치가 증가함에 따라 IMEX가 명시적 방법 대비 달성 가능한 최대 성능 향상은 어느 정도인가?
  • RQ4약 1,000개의 GPU 및 KNL 노드에서 약한 및 강한 확장성 환경에서 IMEX의 성능는 어떻게 확장되는가?
  • RQ5고차수 조절자 및 직접 해법기가 IMEX 스킴에서 암시적 해법의 효율성을 어느 정도 향상시키는가?

주요 결과

  • 쿠르랑트 수치 C = 15일 때, GPU에서 3D-IMEX는 C = 1인 명시적 방법 대비 4배 빠른 성능 확보.
  • 직접 해법기를 사용한 1D-IMEX를 통해 C = 150일 때 최대 100배의 성능 향상 기록, 고쿠르랑트 IMEX 방법의 잠재력 입증.
  • 직접 해법을 사용한 1D-IMEX 스킴은 암시적 해법 중 상호 처리기 간 통신이 감소함에 따라 3D-IMEX보다 뛰어난 확장성 확보.
  • 최대 4,096개 GPU에서의 강한 확장성 분석 결과, 로그-로그 스케일에서 선형 확장성 확보, 명시적 방법 대비 5배의 상대적 성능 향상.
  • 16개 KNL 노드에서 IMEX는 강한 확장성 효율 90%를 달성하여 인텔의 다중코어 아키텍처에서 뛰어난 성능 입증.
  • 글로벌 음파 파동 벤치마크 검증 결과, IMEX 결과는 정확도 면에서 명시적 방법과 일치하지만 성능 면에서 크게 뛰어남.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.