Skip to main content
QUICK REVIEW

[논문 리뷰] Comprehensive Evaluation of OpenCL-based Convolutional Neural Network Accelerators in Xilinx and Altera FPGAs

Ricardo Tapiador, Antonio Ríos-Navarro|arXiv (Cornell University)|2016. 09. 29.
Advanced Neural Network Applications참고 문헌 12인용 수 9
한 줄 요약

이 논문은 5층의 컨볼루션 신경망(CNN)에 대해 Xilinx 및 Altera FPGAs에서 OpenCL 기반의 컨볼루션 신경망(CNN) 가속기 성능을 평가하며, 하드웨어 자원 사용량, 합성 시간, 실행 성능 및 설계 이식성에 대해 비교한다. Xilinx는 더 빠른 합성 시간과 더 나은 자원 활용도를 제공하며, Altera에서는 더 낮은 실행 시간, 다중 플랫폼 지원 및 성숙한 도구 체인으로 인해 성능 면에서 열세를 보인다. 이는 FPGA 기반 딥러닝 가속화에서 효율성과 성능 간의 상충 관계를 보여준다.

ABSTRACT

Deep learning has significantly advanced the state of the art in artificial intelligence, gaining wide popularity from both industry and academia. Special interest is around Convolutional Neural Networks (CNN), which take inspiration from the hierarchical structure of the visual cortex, to form deep layers of convolutional operations, along with fully connected classifiers. Hardware implementations of these deep CNN architectures are challenged with memory bottlenecks that require many convolution and fully-connected layers demanding large amount of communication for parallel computation. Multi-core CPU based solutions have demonstrated their inadequacy for this problem due to the memory wall and low parallelism. Many-core GPU architectures show superior performance but they consume high power and also have memory constraints due to inconsistencies between cache and main memory. FPGA design solutions are also actively being explored, which allow implementing the memory hierarchy using embedded BlockRAM. This boosts the parallel use of shared memory elements between multiple processing units, avoiding data replicability and inconsistencies. This makes FPGAs potentially powerful solutions for real-time classification of CNNs. Both Altera and Xilinx have adopted OpenCL co-design framework from GPU for FPGA designs as a pseudo-automatic development solution. In this paper, a comprehensive evaluation and comparison of Altera and Xilinx OpenCL frameworks for a 5-layer deep CNN is presented. Hardware resources, temporal performance and the OpenCL architecture for CNNs are discussed. Xilinx demonstrates faster synthesis, better FPGA resource utilization and more compact boards. Altera provides multi-platforms tools, mature design community and better execution times.

연구 동기 및 목표

  • 실시간 추론을 위한 Xilinx 및 Altera FPGAs에서 OpenCL 기반의 CNN 가속기 평가 및 비교.
  • 두 플랫폼 간 하드웨어 자원 활용도, 합성 시간 및 실행 성능 분석.
  • FPGA 기반 딥러닝 워크로드에 적합한 OpenCL 도구 체인의 성숙도 및 실용성 평가.
  • 설계 생산성, 성능 및 자원 효율성 측면에서 플랫폼 특화된 이점 식별.

제안 방법

  • Xilinx 및 Altera FPGAs에서 OpenCL을 사용해 5층의 컨볼루션 신경망을 구현.
  • FPGA 내통의 임베디드 BlockRAM을 활용해 메모리 계층을 최적화하고 데이터 이동을 줄임.
  • 여러 FPGA 보드에서 합성 시간, 자원 활용도(LUTs, BRAMs, DSPs), 실행 지연 시간 평가.
  • Xilinx 및 Altera에서 제공하는 OpenCL 프레임워크 비교 분석을 통해 코드 이식성 및 도구 체인의 성숙도에 중점을 둠.
  • 동일한 네트워크 아키텍처와 입력 데이터 조건 하에서 추론 성능 벤치마킹.
  • 다중 프로세서 FPGA 구성에서의 메모리 액세스 패턴 및 병렬 처리 효율성 분석.

실험 결과

연구 질문

  • RQ1Xilinx 및 Altera OpenCL 프레임워크는 CNN 가속화에 있어 합성 시간과 자원 활용도 측면에서 어떻게 비교되는가?
  • RQ25층의 CNN에 대해 Xilinx 및 Altera FPGAs 간 추론 실행 시간에 어떤 성능 차이가 있는가?
  • RQ3두 플랫폼은 설계 생산성, 도구 체인의 성숙도 및 가용한 개발 자원 측면에서 어떻게 다를까?
  • RQ4FPGA 최적화된 메모리 계층은 두 플랫폼 모두에서 CNN 추론 효율성을 얼마나 향상시키는가?
  • RQ5실시간 CNN 추론을 위한 성능, 전력 효율성 및 하드웨어 자원 사용량 간의 최적 트레이드오프를 제공하는 플랫폼은 무엇인가?

주요 결과

  • Xilinx는 더 빠른 합성 시간과 더 나은 FPGA 자원 활용도를 보여, 더 컴act하고 효율적인 하드웨어 구현을 가능하게 했다.
  • Altera에서는 동일한 CNN 워크로드에서 더 낮은 실행 시간을 기록해, 추론 처리량 측면에서 뛰어난 성능을 보였다.
  • Altera는 더 성숙한 도구 체인과 광범위한 다중 플랫폼 지원을 제공해 개발의 유연성과 확장성을 높였다.
  • 두 플랫폼 모두 임베디드 BlockRAM을 효과적으로 활용해 메모리 병목 현상을 줄이고 CNN 처리의 병렬성 향상을 이뤘다.
  • Xilinx 보드는 더 컴act한 크기로 동일한 계산 부하에 대해 더 나은 면적 효율성을 보였다.
  • OpenCL 기반 접근 방식은 두 플랫폼 간 이식 가능한 코드를 가능하게 했지만, 기반 아키텍처와 도구 체인의 차이로 인해 성능에 상당한 격차가 존재했다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.