Skip to main content
QUICK REVIEW

[논문 리뷰] Distributed simulation of polychronous and plastic spiking neural networks: strong and weak scaling of a representative mini-application benchmark executed on a small-scale commodity cluster

Pier Stanislao Paolucci, Roberto Ammendola|arXiv (Cornell University)|2013. 10. 31.
Advanced Memory and Neural Computing참고 문헌 19인용 수 7
한 줄 요약

이 논문은 일반 클러스터에서 다성분 동역학을 가진 유연성 있는 스파iking 신경망을 시뮬레이션하기 위한 천연 분산형 미니애플리케이션 벤치마크인 DPSNN-STDP를 제시한다. 이는 128개의 코어에서 강한 스케일링과 약한 스케일링을 달성하며, 네트워크 활동 1초당 10초의 월클록 시간에 32억 개의 시냅스를 시뮬레이션한다—신경과학 및 HPC 워크로드에 대해 높은 성능와 확장성을 입증한다.

ABSTRACT

We introduce a natively distributed mini-application benchmark representative of plastic spiking neural network simulators. It can be used to measure performances of existing computing platforms and to drive the development of future parallel/distributed computing systems dedicated to the simulation of plastic spiking networks. The mini-application is designed to generate spiking behaviors and synaptic connectivity that do not change when the number of hardware processing nodes is varied, simplifying the quantitative study of scalability on commodity and custom architectures. Here, we present the strong and weak scaling and the profiling of the computational/communication components of the DPSNN-STDP benchmark (Distributed Simulation of Polychronous Spiking Neural Network with synaptic Spike-Timing Dependent Plasticity). In this first test, we used the benchmark to exercise a small-scale cluster of commodity processors (varying the number of used physical cores from 1 to 128). The cluster was interconnected through a commodity network. Bidimensional grids of columns composed of Izhikevich neurons projected synapses locally and toward first, second and third neighboring columns. The size of the simulated network varied from 6.6 Giga synapses down to 200 K synapses. The code demonstrated to be fast and scalable: 10 wall clock seconds were required to simulate one second of activity and plasticity (per Hertz of average firing rate) of a network composed by 3.2 G synapses running on 128 hardware cores clocked @ 2.4 GHz. The mini-application has been designed to be easily interfaced with standard and custom software and hardware communication interfaces. It has been designed from its foundation to be natively distributed and parallel, and should not pose major obstacles against distribution and parallelization on several platforms.

연구 동기 및 목표

  • 탄성 스파이킹 신경망을 시뮬레이션하기 위한 대표성 있고 천연 분산형 미니애플리케이션 벤치마크를 개발하기 위해.
  • 다양한 코어 수를 가진 소규모 일반 클러스터에서 강한 스케일링과 약한 스케일링 성능을 평가하기 위해.
  • 분산 신경망 시뮬레이션에서 계산 및 통신 오버헤드를 정량적으로 평가하기 위해.
  • 향후 대규모 신경망 시뮬레이션에 특화된 병렬/분산 시스템 개발을 지원하기 위해.

제안 방법

  • 벤치마크는 이즈키에비치 뉴런의 네트워크를 이차원 격자로 배열하고, 국소 및 이웃 열기둥 간의 연결(최대 제3 이웃까지)을 구현한다.
  • 시냅스 유연성은 스파이크 타이밍 의존성 유연성(STDP)을 통해 구현되어 시뮬레이션 중에 연결성이 동적으로 변화할 수 있도록 한다.
  • 시뮬레이션은 본질적으로 분산되어 있으며, 데이터와 계산이 처리 노드 간에 분할되어 통신 병목 현상을 최소화하도록 설계되어 있다.
  • 벤치마크는 강한 스케일링(고정된 문제 크기, 변동하는 코어 수)과 약한 스케일링(코어 수에 비례해 문제 크기를 확장)을 모두 지원하여 성능 특성 분석이 가능하다.
  • 성능 병목 현상을 규명하고 로드 밸런스를 평가하기 위해 통신 및 계산 구성 요소를 상세히 프로파일링한다.
  • 표준 네트워크 인터커넥트를 사용하는 일반 클러스터에서 시스템을 실행하며, 맞춤형 하드웨어 의존성이 없다.

실험 결과

연구 질문

  • RQ1DPSNN-STDP 벤치마크는 1에서 128개의 하드웨어 코어에 대해 강한 스케일링과 약한 스케일링에서 어떻게 확장되는가?
  • RQ2분산된 탄성 스파이킹 신경망 시뮬레이션에서 계산 및 통신 오버헤드의 분포는 어떻게 되는가?
  • RQ3대표성 있는 미니애플리케이션 벤치마크는 다양한 수의 처리 노드에서 일관된 스파이크 및 시냅스 행동을 유지할 수 있는가?
  • RQ4일반 클러스터는 현실적인 연결성과 동역학을 가진 대규모 탄성 스파이킹 네트워크를 얼마나 효율적으로 시뮬레이션할 수 있는가?

주요 결과

  • DPSNN-STDP 벤치마크는 128개의 코어에서 32억 개의 시냅스 네트워크를 대상으로 네트워크 활동 1초당 10초의 월클록 시간을 달성했다.
  • 다양한 수의 처리 노드에서 일관된 스파이크 행동과 시냅스 연결성이 유지되어 벤치마크의 이식성과 결정론성을 검증했다.
  • 강한 스케일링에서는 128개의 코어까지 거의 선형적인 성능 향상을 보였으며, 효율적인 로드 분배와 낮은 통신 오버헤드를 시사했다.
  • 약한 스케일링에서는 문제 크기가 코어 수에 비례해 증가함에 따라 높은 효율성을 유지하여, 벤치마크가 대규모 시뮬레이션에 적합함을 확인했다.
  • 프로파일링 결과 계산이 통신보다 지배적이며, 통신 비용은 선형 이하로 증가함을 확인하여 설계의 확장성에 기여했다.
  • 벤치마크는 매우 이식성이 높으며, 주요 수정 없이 표준 및 맞춤형 통신 인터페이스와 통합될 수 있다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.