Skip to main content
QUICK REVIEW

[论文解读] Nonhuman Primate Reaching with Multichannel Sensorimotor Cortex Electrophysiology

O'Doherty, Joseph E., Cardoso, Mariana M. B.|arXiv (Cornell University)|May 26, 2017
Advanced Statistical Methods and Models参考文献 6被引用 4
一句话总结

本文提出了UoI-LASSO和UoI-VAR的可扩展、高性能实现,用于大规模神经生理数据分析,基于超算上的MPI并行化。其达到了迄今最大的向量自回归(VAR)模型(1,000个节点),并在最多278,528个核心上实现了强可扩展性和弱可扩展性,实现了对非人灵长类动物抓握行为数据的可解释且预测性强的建模。

ABSTRACT

<strong>General Description.</strong> This dataset consists of: The threshold crossing times of extracellularly and simultaneously recorded spikes, sorted into units (up to five, including a "hash" unit), along with sorted waveform snippets, and, The x,y position of the fingertip of the reaching hand and the x,y position of reaching targets (both sampled at 250 Hz). The behavioral task was to make self-paced reaches to targets arranged in a grid (e.g. 8x8) without gaps or pre-movement delay intervals. One monkey reached with the right arm (recordings made in the left hemisphere); The other reached with the left arm (right hemisphere). In some sessions recordings were made from both M1 and S1 arrays (192 channels); in most sessions M1 recordings were made alone (96 channels). Data from two primate subjects are included: 37 sessions from monkey 1 ("Indy", spanning about 10 months) and 10 sessions from monkey 2 ("Loco", spanning about 1 month), for a total of ~ 20,000 reaches and 6,500 reaches from monkeys 1 and 2, respectively. <strong>Possible uses. </strong>These data are ideal for training BCI decoders, in particular because they are not segmented into trials. We expect that the dataset will be valuable for researchers who wish to design improved models of sensorimotor cortical spiking or provide an equal footing for comparing different BCI decoders. Other uses could include analyses of the statistics of arm kinematics, spike noise-correlations or signal-correlations, or for exploring the stability or variability of extracellular recording over sessions. <strong>Variable names. </strong>Each file contains data in the following format. In the below, <em>n</em> refers to the number of recording channels, <em>u</em> refers to the number of sorted units, and <em>k</em> refers to the number of samples. chan_names - n x 1 A cell array of channel identifier strings, e.g. "<em>M1 001</em>". cursor_pos - k x 2 The position of the cursor in Cartesian coordinates (x, y), mm. finger_pos - k x 3 <em>or </em>k x 6 The position of the working fingertip in Cartesian coordinates (z, -x, -y), as reported by the hand tracker in cm. Thus the cursor position is an affine transformation of fingertip position using the following matrix:<br> \(\begin{pmatrix} 0 &amp; 0 \\ -10 &amp; 0 \\ 0 &amp; -10 \end{pmatrix}\)<br> Note that for some sessions finger_pos includes the orientation of the sensor as well; the full state is thus: (z, -x, -y, azimuth, elevation, roll). target_pos - k x 2 The position of the target in Cartesian coordinates (x, y), mm. t - k x 1 The timestamp corresponding to each sample of the cursor_pos, finger_pos, and target_pos, seconds. spikes - n x u A cell array of spike event vectors. Each element in the cell array is a vector of spike event timestamps, in seconds. The first unit (<em>u</em><sub>1</sub>) is the "unsorted" unit, meaning it contains the threshold crossings which remained after the spikes on that channel were sorted into other units (<em>u</em><sub>2</sub>, <em>u</em><sub>3</sub>, etc.) For some sessions spikes were sorted into up to 2 units (i.e. <em>u</em>=3); for others, 4 units (<em>u</em>=5). wf - n x u A cell array of spike event waveform "snippets". Each element in the cell array is a matrix of spike event waveforms. Each waveform corresponds to a timestamp in "spikes". Waveform samples are in microvolts. <strong>Videos. </strong>For some sessions, we recorded screencasts of the stimulus presentation display using a dedicated hardware video grabber. These screencasts are thus a faithful representation of the stimuli and feedback presented to the monkey and are available for the following sessions: indy_20160921_01 indy_20160930_02 indy_20160930_05 indy_20161005_06 indy_20161006_02 indy_20161007_02 indy_20161011_03 indy_20161013_03 indy_20161014_04 indy_20161017_02 <strong>Supplements. </strong>The raw broadband neural recordings that the spike trains in this dataset were extracted from are available for the following sessions: indy_20160622_01: doi:10.5281/zenodo.1488440 indy_20160624_03: doi:10.5281/zenodo.1486147 indy_20160627_01: doi:10.5281/zenodo.1484824 indy_20160630_01: doi:10.5281/zenodo.1473703 indy_20160915_01: doi:10.5281/zenodo.1467953 indy_20160916_01: doi:10.5281/zenodo.1467050 indy_20160921_01: doi:10.5281/zenodo.1451793 indy_20160927_04: doi:10.5281/zenodo.1433942 indy_20160927_06: doi:10.5281/zenodo.1432818 indy_20160930_02: doi:10.5281/zenodo.1421880 indy_20160930_05: doi:10.5281/zenodo.1421310 indy_20161005_06: doi:10.5281/zenodo.1419774 indy_20161006_02: doi:10.5281/zenodo.1419172 indy_20161007_02: doi:10.5281/zenodo.1413592 indy_20161011_03: doi:10.5281/zenodo.1412635 indy_20161013_03: doi:10.5281/zenodo.1412094 indy_20161014_04: doi:10.5281/zenodo.1411978 indy_20161017_02: doi:10.5281/zenodo.1411882 indy_20161024_03: doi:10.5281/zenodo.1411474 indy_20161025_04: doi:10.5281/zenodo.1410423 indy_20161026_03: doi:10.5281/zenodo.1321264 indy_20161027_03: doi:10.5281/zenodo.1321256 indy_20161206_02: doi:10.5281/zenodo.1303720 indy_20161207_02: doi:10.5281/zenodo.1302866 indy_20161212_02: doi:10.5281/zenodo.1302832 indy_20161220_02: doi:10.5281/zenodo.1301045 indy_20170123_02: doi:10.5281/zenodo.1167965 indy_20170124_01: doi:10.5281/zenodo.1163026 indy_20170127_03: doi:10.5281/zenodo.1161225 indy_20170131_02: doi:10.5281/zenodo.854733 <strong>Contact Information.</strong> We would be delighted to hear from you if you find this dataset valuable, especially if it leads to publication. Corresponding author: J. E. O'Doherty &lt;joeyo@neuroengineer.com&gt;. <strong>Publications making use of this dataset.</strong> Makin, J. G., O'Doherty, J. E., Cardoso, M. M. B. &amp; Sabes, P. N. (2018). Superior arm-movement decoding from cortex with a new, unsupervised-learning algorithm. <em>J Neural Eng</em> 15(2): 026010. doi:10.1088/1741-2552/aa9e95 Ahmadi, N., Constandinou, T. G., &amp; Bouganis, C.-S. (2018). Spike Rate Estimation Using Bayesian Adaptive Kernel Smoother (BAKS) and Its Application to Brain Machine Interfaces. <em>2018 40th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</em>, Honolulu, HI, USA, 2018, pp. 2547-2550. doi:10.1109/EMBC.2018.8512830 Balasubramanian, M., Ruiz, T., Cook, B., Bhattacharyya, S., Prabhat, Shrivastava, A. &amp; Bouchard K. (2018). Optimizing the Union of Intersections LASSO (UoILASSO) and Vector Autoregressive (UoIVAR) Algorithms for Improved Statistical Estimation at Scale. <em>arXiv preprint arXiv:1808.06992</em> Ahmadi, N., Constandinou, T. G., &amp; Bouganis, C.-S. (2019). Decoding Hand Kinematics from Local Field Potentials Using Long Short-Term Memory (LSTM) Network. <em>arXiv preprint arXiv:1901.00708</em> Clark, D. G., Livezey, J. A., &amp; Bouchard, K. E. (2019). Unsupervised Discovery of Temporal Structure in Noisy Data with Dynamical Components Analysis. <em>arXiv preprint arXiv:1905.09944</em> Shaikh, S., So, R., Sibindi, T., Libedinsky, C., &amp; Basu, A. (2019). Towards Intelligent Intra-cortical BMI (i<sup>2</sup>BMI): Low-power Neuromorphic Decoders that outperform Kalman Filters. <em>bioRxiv preprint</em> 772988. doi:10.1101/772988 Clark, D. G., Livezey, J. A., &amp; Bouchard, K. E. (2019). Unsupervised Discovery of Temporal Structure in Noisy Data with Dynamical Components Analysis. <em>Advances in Neural Information Processing Systems (NeurIPS) 32.</em> Keshtkaran, M. R., &amp; Pandarinath, C. (2019). Enabling hyperparameter optimization in sequential autoencoders for spiking neural data. <em>Advances in Neural Information Processing Systems (NeurIPS) 32.</em> Ahmadi, N., Constandinou, T. G., &amp; Bouganis, C.-S. (2019). End-to-End Hand Kinematic Decoding from LFPs Using Temporal Convolutional Network. <em>2019 IEEE Biomedical Circuits and Systems Conference (BioCAS), </em>Nara, Japan, pp. 1-4. doi:10.1109/biocas.2019.8919131 Bose, S. K., Acharya, J., &amp; Basu, A. (2019). Is my Neural Network Neuromorphic? Taxonomy, Recent Trends and Future Directions in Neuromorphic Engineering. <em>2019 53rd Asilomar Conference on Signals, Systems, and Computers</em>, Pacific Grove, CA, USA, pp. 1522-1527. doi:10.1109/IEEECONF44664.2019.9048891 Sachdeva, P. S., Bhattacharyya, S., &amp; Bouchard, K. E. (2019). Sparse, Predictive, and Interpretable Functional Connectomics with UoILasso, <em>41st Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)</em>, Berlin, Germany, pp. 1965-1968. doi:10.1109/EMBC.2019.8856316 Sachdeva, P. S, Livezey, J. A, Dougherty, M. E., Gu, B.-M., Berke, J. D, &amp; Bouchard, K. E. (2020). Accurate Inference in Parametric Models Reshapes Neuroscientific Interpretation and Improves Data-driven Discovery. <em>bioRxiv Preprint.</em> 2020.04.10.036244. doi:10.1101/2020.04.10.036244 Ahmadi, N., Constandinou, T. G., Bouganis, C.-S. (2020). Inferring entire spiking activity from local field potentials with deep learning. <em>bioRxiv Preprint.</em> 2020.05.02.074104. doi:10.1101/2020.05.02.074104 Ahmadi, N., Constandinou, T. G., Bouganis. C.-S. (2020). Impact of referencing scheme on decoding performance of LFP-based brain-machine interface. <em>bioRxiv Preprint. </em>2020.05.03.075218 doi:10.1101/2020.05.03.075218

研究动机与目标

  • 应对大规模科学数据分析中对可解释且可预测的统计机器学习方法日益增长的需求。
  • 针对高性能计算系统(Cori KNL)上的单节点和多节点执行,优化UoI-LASSO和UoI-VAR算法。
  • 实现对非人灵长类动物抓握任务中高维时间序列神经生理数据的可扩展分析。
  • 在数万个核心上实现强可扩展性和弱可扩展性,以处理TB量级的数据集。
  • 利用真实神经生理数据,在神经科学中构建迄今最大的向量自回归(VAR)模型。

提出的方法

  • 使用C++实现UoI-LASSO,采用基于MPI的并行化,对HDF5文件应用随机数据分发以实现高效的自助抽样。
  • 采用两阶段框架:通过自助样本和正则化参数下LASSO支持的交集进行模型选择,随后通过自助样本上OLS估计的平均值进行模型估计。
  • 使用分布式克罗内克积和向量化方法实现UoI-VAR,以处理高维时间序列数据,并在自助样本和正则化参数上实现并行化。
  • 应用P_B和P_λ并行化策略以缓解通信瓶颈,尤其是在模型选择阶段的MPI_Allreduce调用中。
  • 利用Sparse Eigen C++中的稀疏矩阵运算,优化VAR估计中的矩阵-向量乘法。
  • 采用混合数据分发策略以减少通信开销,特别是在大规模分布式矩阵构建中。

实验结果

研究问题

  • RQ1UoI-LASSO和UoI-VAR框架能否在数万个核心上高效扩展,同时在特征选择和估计中保持低偏差和低方差?
  • RQ2通信开销(尤其是MPI_Allreduce)在大规模运行中对UoI-LASSO和UoI-VAR性能有何影响?
  • RQ3对于大规模数据集,UoI-LASSO(计算与通信)和UoI-VAR(分发与计算)的主要瓶颈分别是什么?
  • RQ4UoI-VAR框架能否在大规模上对非人灵长类动物记录的高维神经时间序列数据进行建模?
  • RQ5在真实神经生理数据上,该框架能够估计的最大VAR模型有多大?

主要发现

  • UoI-LASSO实现可在68至278,528个核心上实现强可扩展性和弱可扩展性,大规模运行中通信开销(尤其是MPI_Allreduce)主导了运行时间。
  • UoI-VAR实现可随着核心数量增加而有效扩展,由于高效的稀疏矩阵运算,计算时间表现出接近理想的强可扩展性。
  • 对于大规模数据集(如1.3TB),通信和分发开销主导了运行时间,尤其在UoI-LASSO的模型选择阶段。
  • UoI-VAR框架成功在真实非人灵长类动物抓握数据上估计了192个节点的VAR模型,使用Cori KNL上的81,600个核心,计算时间为96.9秒,通信时间为1,598.72秒。
  • 本研究构建了神经科学中迄今最大的VAR模型——1,000个节点,证明了UoI-VAR框架在大规模神经数据上的可扩展性和可行性。
  • P_B和P_λ并行化策略被识别为缓解通信瓶颈的关键,尤其在UoI-LASSO的模型选择阶段。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。