Skip to main content
QUICK REVIEW

[论文解读] Efficient non-conjugate Gaussian process factor models for spike count data using polynomial approximations

Stephen Keeley, David M. Zoltowski|arXiv (Cornell University)|Jun 7, 2019
Gaussian Processes and Bayesian Inference参考文献 31被引用 6
一句话总结

本文提出多项式近似似然(PAL),一种快速且准确的方法,用于使用二阶正交多项式近似非共轭高斯过程因子模型中的非线性似然项,从而对神经元放电计数数据进行拟合。PAL 实现了边缘似然的闭式评估和快速优化,在速度和收敛性方面优于 BBVI 和 vLGP,同时为变分推断提供了有效的初始化,在模拟和真实神经数据(涵盖伯努利、泊松和负二项分布观测模型)上表现优异。

ABSTRACT

Gaussian Process Factor Analysis (GPFA) has been broadly applied to the problem of identifying smooth, low-dimensional temporal structure underlying large-scale neural recordings. However, spike trains are non-Gaussian, which motivates combining GPFA with discrete observation models for binned spike count data. The drawback to this approach is that GPFA priors are not conjugate to count model likelihoods, which makes inference challenging. Here we address this obstacle by introducing a fast, approximate inference method for non-conjugate GPFA models. Our approach uses orthogonal second-order polynomials to approximate the nonlinear terms in the non-conjugate log-likelihood, resulting in a method we refer to as extit{polynomial approximate log-likelihood} (PAL) estimators. This approximation allows for accurate closed-form evaluation of marginal likelihoods and fast numerical optimization for parameters and hyperparameters. We derive PAL estimators for GPFA models with binomial, Poisson, and negative binomial observations and find the PAL estimation is highly accurate, and achieves faster convergence times compared to existing state-of-the-art inference methods. We also find that PAL hyperparameters can provide sensible initialization for black box variational inference (BBVI), which improves BBVI accuracy. We demonstrate that PAL estimators achieve fast and accurate extraction of latent structure from multi-neuron spike train data.

研究动机与目标

  • 解决将非共轭推断应用于非高斯放电计数数据的高斯过程因子模型时所面临的挑战。
  • 为具有泊松、伯努利和负二项分布似然的模型开发一种计算高效的替代采样方法,以替代 BBVI 等基于采样的推断方法。
  • 通过使用正交多项式近似非线性对数似然项,实现边缘似然的直接优化。
  • 为黑箱变分推断(BBVI)提供稳健的初始参数估计,以提升其稳定性和收敛性。
  • 在小鼠和灵长类皮层的模拟和真实多神经元神经数据集上,展示该方法在准确性和速度方面的表现。

提出的方法

  • 使用二阶正交多项式近似 GPFA 模型中具有计数观测值的非共轭对数似然中的非线性项。
  • 通过在多项式近似下对潜变量进行积分,推导出边缘似然的闭式表达式。
  • 将该方法应用于三种观测模型:伯努利分布、泊松分布和负二项分布,每种模型分别捕捉不同的神经放电分散特性。
  • 采用数值优化直接最大化近似后的边缘似然,避免了随机梯度估计和采样方差。
  • 将 PAL 估计的超参数用作 BBVI 的初始化,从而提升其收敛性和准确性。
  • 通过模拟和来自小鼠视觉皮层及灵长类顶叶皮层的真实神经数据验证该方法。

实验结果

研究问题

  • RQ1多项式近似非共轭对数似然是否能实现在放电计数数据的 GPFA 模型中快速、闭式边缘似然评估?
  • RQ2在模拟和真实神经数据上,PAL 在收敛速度、准确性和稳定性方面与 BBVI 和 vLGP 相比表现如何?
  • RQ3PAL 是否能为 BBVI 提供有效的初始化,从而改善其收敛性并降低超参数估计的方差?
  • RQ4在真实神经数据中,伯努利、泊松或负二项分布模型中哪一种拟合效果最佳?PAL 在每种模型下是否能准确恢复潜在结构?
  • RQ5由 PAL 估计模型恢复的潜在结构是否反映了生物上合理的神经动力学,例如在决策任务中的选择编码?

主要发现

  • PAL 在模拟和真实神经数据上的表现与 BBVI 和 vLGP 相当或更优,且收敛时间显著更短。
  • PAL 提供了闭式边缘似然,使直接优化成为可能,无需调整学习率或蒙特卡洛样本数量。
  • 对于伯努利分布和负二项分布 GPFA 模型,PAL 表现与 BBVI 相当,且在小鼠视觉皮层数据上表现出更高的交叉验证对数似然。
  • 在灵长类顶叶皮层数据集中,由 PAL 初始化的 BBVI 表现优于单独使用 BBVI,尤其在负二项分布模型中,表明其稳定性与准确性均得到提升。
  • 在小鼠数据中,伯努利-GPFA 模型获得了最高的交叉验证对数似然,表明其可能被低估但对某些神经元群体具有有效性。
  • 由伯努利-GPFA 恢复的潜在结构显示出一个维度在刺激后约 400ms 开始发散,与决策过程时间点一致,提示其可能编码了选择变量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。