Skip to main content
QUICK REVIEW

[논문 리뷰] Nonhuman Primate Reaching with Multichannel Sensorimotor Cortex Electrophysiology

O'Doherty, Joseph E., Cardoso, Mariana M. B.|arXiv (Cornell University)|2017. 05. 26.
Advanced Statistical Methods and Models참고 문헌 6인용 수 4
한 줄 요약

이 논문은 슈퍼컴퓨터에서 MPI 기반 병렬 처리를 활용하여 대규모 신경생리학적 데이터 분석을 위한 확장성 있고 고성능의 UoI-LASSO 및 UoI-VAR 구현을 제시한다. 이는 가장 큰 알려진 VAR 모델(1,000개 노드)을 달성했으며, 최대 278,528개의 코어를 활용한 강한 스케일링과 약한 스케일링을 보여주어 비인간 영장류의 움직임 데이터에 대한 해석 가능하고 예측 가능한 모델링을 가능하게 한다.

ABSTRACT

<strong>General Description.</strong> This dataset consists of: The threshold crossing times of extracellularly and simultaneously recorded spikes, sorted into units (up to five, including a "hash" unit), along with sorted waveform snippets, and, The x,y position of the fingertip of the reaching hand and the x,y position of reaching targets (both sampled at 250 Hz). The behavioral task was to make self-paced reaches to targets arranged in a grid (e.g. 8x8) without gaps or pre-movement delay intervals. One monkey reached with the right arm (recordings made in the left hemisphere); The other reached with the left arm (right hemisphere). In some sessions recordings were made from both M1 and S1 arrays (192 channels); in most sessions M1 recordings were made alone (96 channels). Data from two primate subjects are included: 37 sessions from monkey 1 ("Indy", spanning about 10 months) and 10 sessions from monkey 2 ("Loco", spanning about 1 month), for a total of ~ 20,000 reaches and 6,500 reaches from monkeys 1 and 2, respectively. <strong>Possible uses. </strong>These data are ideal for training BCI decoders, in particular because they are not segmented into trials. We expect that the dataset will be valuable for researchers who wish to design improved models of sensorimotor cortical spiking or provide an equal footing for comparing different BCI decoders. Other uses could include analyses of the statistics of arm kinematics, spike noise-correlations or signal-correlations, or for exploring the stability or variability of extracellular recording over sessions. <strong>Variable names. </strong>Each file contains data in the following format. In the below, <em>n</em> refers to the number of recording channels, <em>u</em> refers to the number of sorted units, and <em>k</em> refers to the number of samples. chan_names - n x 1 A cell array of channel identifier strings, e.g. "<em>M1 001</em>". cursor_pos - k x 2 The position of the cursor in Cartesian coordinates (x, y), mm. finger_pos - k x 3 <em>or </em>k x 6 The position of the working fingertip in Cartesian coordinates (z, -x, -y), as reported by the hand tracker in cm. Thus the cursor position is an affine transformation of fingertip position using the following matrix:<br> \(\begin{pmatrix} 0 &amp; 0 \\ -10 &amp; 0 \\ 0 &amp; -10 \end{pmatrix}\)<br> Note that for some sessions finger_pos includes the orientation of the sensor as well; the full state is thus: (z, -x, -y, azimuth, elevation, roll). target_pos - k x 2 The position of the target in Cartesian coordinates (x, y), mm. t - k x 1 The timestamp corresponding to each sample of the cursor_pos, finger_pos, and target_pos, seconds. spikes - n x u A cell array of spike event vectors. Each element in the cell array is a vector of spike event timestamps, in seconds. The first unit (<em>u</em><sub>1</sub>) is the "unsorted" unit, meaning it contains the threshold crossings which remained after the spikes on that channel were sorted into other units (<em>u</em><sub>2</sub>, <em>u</em><sub>3</sub>, etc.) For some sessions spikes were sorted into up to 2 units (i.e. <em>u</em>=3); for others, 4 units (<em>u</em>=5). wf - n x u A cell array of spike event waveform "snippets". Each element in the cell array is a matrix of spike event waveforms. Each waveform corresponds to a timestamp in "spikes". Waveform samples are in microvolts. <strong>Videos. </strong>For some sessions, we recorded screencasts of the stimulus presentation display using a dedicated hardware video grabber. These screencasts are thus a faithful representation of the stimuli and feedback presented to the monkey and are available for the following sessions: indy_20160921_01 indy_20160930_02 indy_20160930_05 indy_20161005_06 indy_20161006_02 indy_20161007_02 indy_20161011_03 indy_20161013_03 indy_20161014_04 indy_20161017_02 <strong>Supplements. </strong>The raw broadband neural recordings that the spike trains in this dataset were extracted from are available for the following sessions: indy_20160622_01: doi:10.5281/zenodo.1488440 indy_20160624_03: doi:10.5281/zenodo.1486147 indy_20160627_01: doi:10.5281/zenodo.1484824 indy_20160630_01: doi:10.5281/zenodo.1473703 indy_20160915_01: doi:10.5281/zenodo.1467953 indy_20160916_01: doi:10.5281/zenodo.1467050 indy_20160921_01: doi:10.5281/zenodo.1451793 indy_20160927_04: doi:10.5281/zenodo.1433942 indy_20160927_06: doi:10.5281/zenodo.1432818 indy_20160930_02: doi:10.5281/zenodo.1421880 indy_20160930_05: doi:10.5281/zenodo.1421310 indy_20161005_06: doi:10.5281/zenodo.1419774 indy_20161006_02: doi:10.5281/zenodo.1419172 indy_20161007_02: doi:10.5281/zenodo.1413592 indy_20161011_03: doi:10.5281/zenodo.1412635 indy_20161013_03: doi:10.5281/zenodo.1412094 indy_20161014_04: doi:10.5281/zenodo.1411978 indy_20161017_02: doi:10.5281/zenodo.1411882 indy_20161024_03: doi:10.5281/zenodo.1411474 indy_20161025_04: doi:10.5281/zenodo.1410423 indy_20161026_03: doi:10.5281/zenodo.1321264 indy_20161027_03: doi:10.5281/zenodo.1321256 indy_20161206_02: doi:10.5281/zenodo.1303720 indy_20161207_02: doi:10.5281/zenodo.1302866 indy_20161212_02: doi:10.5281/zenodo.1302832 indy_20161220_02: doi:10.5281/zenodo.1301045 indy_20170123_02: doi:10.5281/zenodo.1167965 indy_20170124_01: doi:10.5281/zenodo.1163026 indy_20170127_03: doi:10.5281/zenodo.1161225 indy_20170131_02: doi:10.5281/zenodo.854733 <strong>Contact Information.</strong> We would be delighted to hear from you if you find this dataset valuable, especially if it leads to publication. Corresponding author: J. E. O'Doherty &lt;joeyo@neuroengineer.com&gt;. <strong>Publications making use of this dataset.</strong> Makin, J. G., O'Doherty, J. E., Cardoso, M. M. B. &amp; Sabes, P. N. (2018). Superior arm-movement decoding from cortex with a new, unsupervised-learning algorithm. <em>J Neural Eng</em> 15(2): 026010. doi:10.1088/1741-2552/aa9e95 Ahmadi, N., Constandinou, T. G., &amp; Bouganis, C.-S. (2018). Spike Rate Estimation Using Bayesian Adaptive Kernel Smoother (BAKS) and Its Application to Brain Machine Interfaces. <em>2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</em>, Honolulu, HI, USA, 2018, pp. 2547-2550. doi:10.1109/EMBC.2018.8512830 Balasubramanian, M., Ruiz, T., Cook, B., Bhattacharyya, S., Prabhat, Shrivastava, A. &amp; Bouchard K. (2018). Optimizing the Union of Intersections LASSO (UoILASSO) and Vector Autoregressive (UoIVAR) Algorithms for Improved Statistical Estimation at Scale. <em>arXiv preprint arXiv:1808.06992</em> Ahmadi, N., Constandinou, T. G., &amp; Bouganis, C.-S. (2019). Decoding Hand Kinematics from Local Field Potentials Using Long Short-Term Memory (LSTM) Network. <em>arXiv preprint arXiv:1901.00708</em> Clark, D. G., Livezey, J. A., &amp; Bouchard, K. E. (2019). Unsupervised Discovery of Temporal Structure in Noisy Data with Dynamical Components Analysis. <em>arXiv preprint arXiv:1905.09944</em> Shaikh, S., So, R., Sibindi, T., Libedinsky, C., &amp; Basu, A. (2019). Towards Intelligent Intra-cortical BMI (i<sup>2</sup>BMI): Low-power Neuromorphic Decoders that outperform Kalman Filters. <em>bioRxiv preprint</em> 772988. doi:10.1101/772988 Clark, D. G., Livezey, J. A., &amp; Bouchard, K. E. (2019). Unsupervised Discovery of Temporal Structure in Noisy Data with Dynamical Components Analysis. <em>Advances in Neural Information Processing Systems (NeurIPS) 32.</em> Keshtkaran, M. R., &amp; Pandarinath, C. (2019). Enabling hyperparameter optimization in sequential autoencoders for spiking neural data. <em>Advances in Neural Information Processing Systems (NeurIPS) 32.</em> Ahmadi, N., Constandinou, T. G., &amp; Bouganis, C.-S. (2019). End-to-End Hand Kinematic Decoding from LFPs Using Temporal Convolutional Network. <em>2019 IEEE Biomedical Circuits and Systems Conference (BioCAS), </em>Nara, Japan, pp. 1-4. doi:10.1109/biocas.2019.8919131 Bose, S. K., Acharya, J., &amp; Basu, A. (2019). Is my Neural Network Neuromorphic? Taxonomy, Recent Trends and Future Directions in Neuromorphic Engineering. <em>2019 53rd Asilomar Conference on Signals, Systems, and Computers</em>, Pacific Grove, CA, USA, pp. 1522-1527. doi:10.1109/IEEECONF44664.2019.9048891 Sachdeva, P. S., Bhattacharyya, S., &amp; Bouchard, K. E. (2019). Sparse, Predictive, and Interpretable Functional Connectomics with UoILasso, <em>41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</em>, Berlin, Germany, pp. 1965-1968. doi:10.1109/EMBC.2019.8856316 Sachdeva, P. S, Livezey, J. A, Dougherty, M. E., Gu, B.-M., Berke, J. D, &amp; Bouchard, K. E. (2020). Accurate Inference in Parametric Models Reshapes Neuroscientific Interpretation and Improves Data-driven Discovery. <em>bioRxiv Preprint.</em> 2020.04.10.036244. doi:10.1101/2020.04.10.036244 Ahmadi, N., Constandinou, T. G., Bouganis, C.-S. (2020). Inferring entire spiking activity from local field potentials with deep learning. <em>bioRxiv Preprint.</em> 2020.05.02.074104. doi:10.1101/2020.05.02.074104 Ahmadi, N., Constandinou, T. G., Bouganis. C.-S. (2020). Impact of referencing scheme on decoding performance of LFP-based brain-machine interface. <em>bioRxiv Preprint. </em>2020.05.03.075218 doi:10.1101/2020.05.03.075218

연구 동기 및 목표

  • 대규모 과학적 데이터 분석에서 해석 가능하고 예측 가능한 통계적 기계학습 방법의 증가하는 수요를 충족시키기 위해.
  • 고성능 컴퓨팅 시스템(Cori KNL)에서 단일 노드 및 다중 노드 실행을 위한 UoI-LASSO 및 UoI-VAR 알고리즘 최적화를 위해.
  • 비인간 영장류의 움직임 작업에서 유래한 고차원 시계열 신경생리학적 데이터의 확장 가능한 분석을 위해.
  • 테라바이트 규모의 데이터셋을 처리하기 위해 수만 개의 코어를 활용한 강한 스케일링과 약한 스케일링을 달성하기 위해.
  • 실제 신경생리학적 데이터를 사용하여 신경과학 분야에서 가장 큰 벡터 자기회귀(VAR) 모델을 만들기 위해.

제안 방법

  • MPI 기반 병렬 처리를 사용한 C++ 기반 UoI-LASSO의 구현으로, 효율적인 부트스트랩 표본 추출을 위해 HDF5 파일에 랜덤 데이터 분포를 적용한다.
  • 이중 단계 프레임워크를 사용: 첫 번째 단계는 부트스트랩 표본과 정규화 파라미터를 기반으로 한 LASSO 지원 집합의 교차를 통한 모델 선택이며, 두 번째 단계는 부트스트랩 표본 간의 최소제곱법(OLS) 평균을 통한 모델 추정.
  • 분산 크로네커 곱과 벡터화를 활용하여 고차원 시계열 데이터를 처리하는 UoI-VAR의 구현으로, 부트스트랩 표본과 정규화 파라미터에 대한 병렬 처리를 수행한다.
  • 통신 병목 현상을 완화하기 위해 P_B 및 P_λ 병렬 처리 전략을 적용하여, 특히 모델 선택 단계에서 발생하는 MPI_Allreduce 호출의 영향을 줄인다.
  • Sparse Eigen C++를 활용해 희소 행렬 연산을 최적화하여 VAR 추정에서 행렬-벡터 곱 연산을 효율적으로 처리한다.
  • 특히 대규모 분산 행렬 구축에서의 통신 오버헤드를 줄이기 위해 하이브리드 데이터 분포 전략을 활용한다.

실험 결과

연구 질문

  • RQ1UoI-LASSO 및 UoI-VAR 프레임워크는 수만 개의 코어에서 실행될 수 있을 정도로 효율적으로 확장될 수 있는가? 이때 특징 선택 및 추정에서 낮은 편향과 낮은 분산을 유지할 수 있는가?
  • RQ2대규모에서의 성능에 영향을 미치는 통신 오버헤드, 특히 MPI_Allreduce에서 비롯된 오버헤드는 UoI-LASSO 및 UoI-VAR에 어떤 영향을 미치는가?
  • RQ3대규모 데이터셋에서 UoI-LASSO(계산 대비 통신) 및 UoI-VAR(분포 대비 계산)의 주요 병목 현상은 무엇인가?
  • RQ4UoI-VAR 프레임워크는 대규모에서 비인간 영장류 기록 데이터로부터 고차원 신경 시계열 데이터를 모델링하는 데 사용될 수 있는가?
  • RQ5이 프레임워크를 사용하여 실제 신경생리학적 데이터에서 추정할 수 있는 가장 큰 VAR 모델은 무엇인가?

주요 결과

  • UoI-LASSO의 구현은 68에서 278,528개의 코어에 걸쳐 강한 스케일링과 약한 스케일링을 달성했으며, 대규모 실행에서 통신(특히 MPI_Allreduce)이 런타임의 주요 요소로 작용했다.
  • UoI-VAR의 구현은 코어 수 증가에 따라 효과적으로 스케일링되었으며, 효율적인 희소 행렬 연산 덕분에 계산 시간이 거의 이상적인 강한 스케일링을 보였다.
  • 대규모 데이터셋(예: 1.3TB)에서는 통신 및 분포 오버헤드가 런타임의 주요 요소로 작용했으며, 특히 UoI-LASSO의 모델 선택 단계에서 두드러졌다.
  • UoI-VAR 프레임워크는 실제로 비인간 영장류의 움직임 데이터에서 192개 노드의 VAR 모델을 성공적으로 추정했으며, Cori KNL에서 81,600개의 코어를 사용해 계산 시간 96.9초, 통신 시간 1,598.72초를 기록했다.
  • 본 연구는 신경과학 분야에서 알려진 바에 따르면 가장 큰 VAR 모델인 1,000개 노드를 구축하여, UoI-VAR 프레임워크의 확장성과 대규모 신경 데이터에 대한 실현 가능성을 입증했다.
  • P_B 및 P_λ 병렬 처리 전략이 특히 UoI-LASSO의 모델 선택 단계에서 통신 병목 현상을 완화하는 데 핵심적인 역할을 한다고 규명되었다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.