[논문 리뷰] Open SYCL on heterogeneous GPU systems: A case of study
이 논문은 유한 시간 리아풀로프 지수 계산을 사례 연구로 삼아 이질적 GPU 시스템에서 Open SYCL을 평가하며, CUDA 및 HIP와 유사한 성능을 제공하면서도 호스트 코드 개발 노력은 크게 줄이고 NVIDIA 및 AMD 장치 간의 원활한 다중 벤더 GPU 활용을 가능하게 함을 보여준다.
Computational platforms for high-performance scientific applications are becoming more heterogenous, including hardware accelerators such as multiple GPUs. Applications in a wide variety of scientific fields require an efficient and careful management of the computational resources of this type of hardware to obtain the best possible performance. However, there are currently different GPU vendors, architectures and families that can be found in heterogeneous clusters or machines. Programming with the vendor provided languages or frameworks, and optimizing for specific devices, may become cumbersome and compromise portability to other systems. To overcome this problem, several proposals for high-level heterogeneous programming have appeared, trying to reduce the development effort and increase functional and performance portability, specifically when using GPU hardware accelerators. This paper evaluates the SYCL programming model, using the Open SYCL compiler, from two different perspectives: The performance it offers when dealing with single or multiple GPU devices from the same or different vendors, and the development effort required to implement the code. We use as case of study the Finite Time Lyapunov Exponent calculation over two real-world scenarios and compare the performance and the development effort of its Open SYCL-based version against the equivalent versions that use CUDA or HIP. Based on the experimental results, we observe that the use of SYCL does not lead to a remarkable overhead in terms of the GPU kernels execution time. In general terms, the Open SYCL development effort for the host code is lower than that observed with CUDA or HIP. Moreover, the SYCL version can take advantage of both CUDA and AMD GPU devices simultaneously much easier than directly using the vendor-specific programming solutions.
연구 동기 및 목표
- 다양한 벤더와 아키텍처를 가진 이질적 GPU 클러스터 프로그래밍의 증가하는 도전 과제를 해결하기 위해.
- CUDA 및 HIP와 같은 벤더 고유의 프로그래밍 모델에서 유발되는 복잡성과 이식성 문제를 줄이기 위해.
- 단일 및 다중 벤더 GPU 배포를 모두 지원하는 고수준이고 이식 가능한 대안으로 Open SYCL을 평가하기 위해.
- 실제 과학 HPC 워크로드에서 성능와 개발 노력 간의 상호 교환 관계를 평가하기 위해.
- 혼합 GPU 환경에서 생산 수준의 과학 응용 프로그램에 SYCL을 사용할 수 있는지 실현 가능성을 입증하기 위해.
제안 방법
- 직접 비교를 위해 Open SYCL, CUDA 및 HIP를 사용하여 유한 시간 리아풀로프 지수(FTLE) 계산을 구현하기 위해.
- Open SYCL 컴파일러를 사용하여 단일 소스 코드에서 NVIDIA 및 AMD GPU 장치를 대상으로 커널을 생성하기 위해.
- 모든 세 구현에서 커널 실행 시간과 호스트 코드 복잡성을 프로파일링하기 위해.
- 혼합 GPU 벤더(NVIDIA 및 AMD)가 포함된 이질적 시스템에서 성능을 측정하여 이식성과 로드 밸런싱을 평가하기 위해.
- 호스트 코드 구조와 코드 라인 수를 분석하여 개발 노력의 차이를 정량화하기 위해.
- 성능 및 생산성의 차이를 고립시키기 위해 모든 구현 간의 기능 일치를 확보하기 위해.
실험 결과
연구 질문
- RQ1Open SYCL은 벤더 최적화된 CUDA 및 HIP 구현에 비해 상당한 성능 오버헤드를 유발하는가?
- RQ2Open SYCL에서 호스트 코드 개발 노력은 CUDA 및 HIP에 비해 어떻게 다른가?
- RQ3SYCL은 동일한 응용 프로그램에서 서로 다른 벤더의 여러 GPU 장치를 얼마나 효율적으로 활용할 수 있는가?
- RQ4SYCL은 효율성을 희생시키지 않고 다양한 GPU 아키텍처 간에 성능 이식성을 달성할 수 있는가?
- RQ5이질적 환경에서 고성능 과학 계산을 위해 Open SYCL을 실용적으로 사용할 수 있는가?
주요 결과
- Open SYCL은 CUDA 및 HIP와 동일한 GPU 커널 실행 시간을 달성하여 유의미한 성능 오버헤드가 관찰되지 않았다.
- 더 높은 수준의 추상화와 줄어든 버일러플레이트 덕분에 Open SYCL에서 호스트 코드 개발 노력은 CUDA 또는 HIP보다 상당히 낮았다.
- SYCL은 이질적 GPU 시스템을 네이티브로 지원하여 최소한의 코드 변경으로도 NVIDIA 및 AMD GPU를 동시에 효율적으로 사용할 수 있었다.
- 벤더 고유 솔루션에 비해 SYCL의 구현이 더 뛰어난 이식성을 보였다.
- 단일 GPU 및 다중 GPU 구성 모두에서 SYCL 기반 구현의 성능가 경쟁력 있는 수준을 유지했다.
- 이 연구는 SYCL이 다중 벤더 GPU 지원과 유지보수 가능한 코드베이스가 필요한 과학 HPC 응용 프로그램을 위한 실현 가능한 대안가 될 수 있음을 확인했다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.