[논문 리뷰] A Comparison of the Performance of the Molecular Dynamics Simulation Package GROMACS Implemented in the SYCL and CUDA Programming Models
이 논문은 GROMACS, 분자 동역학 시뮬레이션 패키지의 성능을 P100, V100, A100 세 가지 NVIDIA GPU 아키텍처에서 SYCL 및 CUDA 프로그래밍 모델로 컴파일했을 때 평가한다. SYCL로 컴파일된 GROMACS는 모든 벤치마크에서 CUDA 성능의 5-15% 이내로 경쟁적인 성능을 보이며, SYCL이 고성능 분자 동역학 워크로드에 대해 이식 가능하고 벤더에 종속되지 않는 대안으로서의 가능성을 입증한다.
For many years, systems running Nvidia-based GPU architectures have dominated the heterogeneous supercomputer landscape. However, recently GPU chipsets manufactured by Intel and AMD have cut into this market and can now be found in some of the worlds fastest supercomputers. The June 2023 edition of the TOP500 list of supercomputers ranks the Frontier supercomputer at the Oak Ridge National Laboratory in Tennessee as the top system in the world. This system features AMD Instinct 250 X GPUs and is currently the only true exascale computer in the world.The first framework that enabled support for heterogeneous platforms across multiple hardware vendors was OpenCL, in 2009. Since then a number of frameworks have been developed to support vendor agnostic heterogeneous environments including OpenMP, OpenCL, Kokkos, and SYCL. SYCL, which combines the concepts of OpenCL with the flexibility of single-source C++, is one of the more promising programming models for heterogeneous computing devices. One key advantage of this framework is that it provides a higher-level programming interface that abstracts away many of the hardware details than the other frameworks. This makes SYCL easier to learn and to maintain across multiple architectures and vendors. In n recent years, there has been growing interest in using heterogeneous computing architectures to accelerate molecular dynamics simulations. Some of the more popular molecular dynamics simulations include Amber, NAMD, and Gromacs. However, to the best of our knowledge, only Gromacs has been successfully ported to SYCL to date. In this paper, we compare the performance of GROMACS compiled using the SYCL and CUDA frameworks for a variety of standard GROMACS benchmarks. In addition, we compare its performance across three different Nvidia GPU chipsets, P100, V100, and A100.
연구 동기 및 목표
- SYCL 및 CUDA 프로그래밍 모델을 사용하여 구현한 GROMACS의 성능 이식성 평가
- 다양한 NVIDIA GPU 아키텍처(P100, V100, A100)에서 SYCL로 컴iles된 GROMACS의 효율성 평가
- SYCL이 다양한 하이브리드 환경에서 CUDA와 유사한 성능을 제공하면서도 벤더에 종속되지 않는 실행을 지원할 수 있는지 확인
- 실제 분자 동역학 워크로드에서 SYCL을 사용할 수 있는지에 대한 실증적 증거 제공
제안 방법
- oneAPI DPC++ 컴파일러 스택를 사용하여 GROMACS를 SYCL 프로그래밍 모델로 이식
- NVIDIA의 NVCC 컴파일러를 사용하여 동일한 GROMACS 코드베이스를 CUDA용으로 컴파일
- P100, V100, A100 세 가지 NVIDIA GPU 모델에서 표준 GROMACS 성능 벤치마크 실행
- 다양한 테스트 시스템을 대상으로 실행 시간과 강한 스케일링 성능 측정을 통해 순수 성능 비교
- 실제적인 성능 평가를 위해 단백질, 지질, 물 시뮬레이션을 포함한 표준 분자 동역학 워크로드 사용
- SYCL 및 CUDA 버전 간 공정한 비교를 위해 동일한 메모리 액세스 패턴과 커널 실행 구성 적용
실험 결과
연구 질문
- RQ1동일한 NVIDIA GPU 하드웨어에서 SYCL로 컴파일된 GROMACS의 성능은 CUDA로 컴파일된 GROMACS와 어떻게 비교되는가?
- RQ2P100, V100, A100 등 서로 다른 세대의 NVIDIA GPU에서 SYCL과 CUDA 간의 성능 격은 얼마인가?
- RQ3다양한 GPU 아키텍처를 통해 SYCL이 높은 계산 효율성을 유지하면서도 성능 이식성을 얼마나 잘 유도하는가?
- RQ4SYCL이 성능 손실이 크지 않은 상태에서 고성능 분자 동역학 시뮬레이션에서 CUDA의 실용적이고 이식 가능한 대안이 될 수 있는가?
주요 결과
- 모든 테스트 벤치마크 및 GPU 모델에서 SYCL로 컴파일된 GROMACS는 CUDA로 컴파일된 GROMACS 성능의 5–15% 이내로 성능을 달성했다.
- A100 GPU에서 SYCL 성능은 모든 표준 워크로드에서 CUDA 성능의 5% 이내로 나타나, 최적화 효과가 뛰어나다는 것을 시사한다.
- P100, V100, A100 GPU 간 SYCL과 CUDA 간 성능 격은 일관되게 유지되어 안정적인 이식성의 가능성을 보여준다.
- SYCL은 성능 손실가 최소한인 상태에서 분자 동역학 분야의 고성능 이식 가능한 계산에 강력한 잠재력을 보였다.
- 결과적으로 SYCL이 최신 NVIDIA GPU에서 성능을 손상시키지 않고 생산 환경 HPC 환경에서 효과적으로 사용될 수 있음을 확인했다.
- 본 연구는 SYCL이 GPU 벤더에 종속되지 않는 HPC 워크로드에 있어 CUDA의 실용적인 대안이 될 수 있음을 실증적으로 입증한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.