[论文解读] tcFFT: Accelerating Half-Precision FFT through Tensor Cores
tcFFT 是一个高性能 FFT 库,通过在 Tensor Core 片段上启用单元素操作并优化内存访问模式,利用 NVIDIA Tensor Cores 加速半精度 1D 和 2D FFT。在 V100 和 A100 GPU 上,tcFFT 相较于 cuFFT 实现了 1.29x–3.24x 的加速,适用于多种 FFT 大小和维度。
Fast Fourier Transform (FFT) is an essential tool in scientific and engineering computation. The increasing demand for mixed-precision FFT has made it possible to utilize half-precision floating-point (FP16) arithmetic for faster speed and energy saving. Specializing in lower precision, NVIDIA Tensor Cores can deliver extremely high computation performance. However, the fixed computation pattern makes it hard to utilize the computing power of Tensor Cores in FFT. Therefore, we developed tcFFT to accelerate FFT with Tensor Cores. Our tcFFT supports batched 1D and 2D FFT of various sizes and it exploits a set of optimizations to achieve high performance: 1) single-element manipulation on Tensor Core fragments to support special operations needed by FFT; 2) fine-grained data arrangement design to coordinate with the GPU memory access pattern. We evaluated our tcFFT and the NVIDIA cuFFT in various sizes and dimensions on NVIDIA V100 and A100 GPUs. The results show that our tcFFT can outperform cuFFT 1.29x-3.24x and 1.10x-3.03x on the two GPUs, respectively. Our tcFFT has a great potential for mixed-precision scientific applications.
研究动机与目标
- 通过利用为低精度 GEMM 优化但并非原生适用于 FFT 特殊操作的 NVIDIA Tensor Cores,加速现代 GPU 上的半精度 FFT。
- 克服由于大尺寸或多维 FFT 中不规则内存访问模式导致的性能瓶颈。
- 设计一个可移植的高性能 FFT 库,支持广泛尺寸下的批量 1D 和 2D 变换。
- 实现对 FFT 工作负载的 Tensor Cores 高效利用,这些工作负载在现有库(如 cuFFT)中支持不佳。
- 证明 Tensor Cores 加速可有效应用于传统密集线性代数之外的科学计算工作负载。
提出的方法
- 提出在 Tensor Core 片段上进行单元素操作的技术,以支持 FFT 特有的操作,如复矩阵访问和逐元素乘法。
- 设计细粒度的数据排列策略,实现在不同 FFT 尺寸下连续、合并的内存访问模式。
- 重新组织全局内存中的数据布局,以最小化大尺寸或 2D FFT 中合并阶段的非合并步长访问。
- 利用共享内存减少全局内存带宽压力并提高计算强度。
- 实施内核融合策略,将 FFT 计算阶段映射到 Tensor Cores 操作,同时最小化数据移动。
- 使用 Tensor Core WMMA 接口,结合自定义内存布局和线程映射,以最大化占用率和利用率。
实验结果
研究问题
- RQ1尽管 Tensor Cores 的设计固定为 GEMM,是否能有效利用它们来加速半精度 FFT?
- RQ2如何缓解大尺寸或多维 FFT 中的不规则内存访问模式,以避免 GPU 内存瓶颈?
- RQ3为在 Tensor Cores 上支持 FFT 的特殊操作(如复数运算和逐元素缩放),需要哪些优化?
- RQ4一个单一的 FFT 库能否通过 Tensor Cores 在广泛范围的 1D 和 2D FFT 尺寸上实现一致的性能提升?
- RQ5与广泛使用的 cuFFT 相比,Tensor Cores 优化的 FFT 库在真实科学计算工作负载中的性能表现如何?
主要发现
- 在 V100 GPU 上,tcFFT 相较于 cuFFT 在 1D FFT 上实现了平均 1.90x 的加速。
- 在 2D FFT 上,tcFFT 在 V100 上相比 cuFFT 加速 1.29x 到 3.24x,在 A100 上加速 1.10x 到 3.03x。
- tcFFT 的性能优势在大 FFT 尺寸和高批量数时最为显著,此时 cuFFT 的内存访问模式显著恶化。
- 即使在小批量尺寸下,tcFFT 仍保持高性能,在 1D FFT 中批量大小超过 4 时优于 cuFFT,在 2D FFT 中批量大小超过 2 时优于 cuFFT。
- 所提出的单元素片段操作技术使 FFT 特有操作在 Tensor Cores 上高效执行,克服了硬件原生支持的关键限制。
- 优化的内存布局确保了连续且合并的全局内存访问,显著减少了大规模 FFT 中的内存带宽瓶颈。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。