[论文解读] Infinitely Wide Tensor Networks as Gaussian Process
本文证明了无限宽张量网络——具体而言是纯矩阵乘积态(MPS)以及两种混合架构(神经核MPS与含隐藏神经层的MPS)——在无限宽极限下收敛于高斯过程(GPs)。通过推导所诱导GP的均值函数与协方差函数,作者表明先验标准差等超参数控制GP的特征长度尺度,数值实验进一步证实,先验方差增大时,样本路径的复杂度也随之提高。
Gaussian Process is a non-parametric prior which can be understood as a distribution on the function space intuitively. It is known that by introducing appropriate prior to the weights of the neural networks, Gaussian Process can be obtained by taking the infinite-width limit of the Bayesian neural networks from a Bayesian perspective. In this paper, we explore the infinitely wide Tensor Networks and show the equivalence of the infinitely wide Tensor Networks and the Gaussian Process. We study the pure Tensor Network and another two extended Tensor Network structures: Neural Kernel Tensor Network and Tensor Network hidden layer Neural Network and prove that each one will converge to the Gaussian Process as the width of each model goes to infinity. (We note here that Gaussian Process can also be obtained by taking the infinite limit of at least one of the bond dimensions $α_{i}$ in the product of tensor nodes, and the proofs can be done with the same ideas in the proofs of the infinite-width cases.) We calculate the mean function (mean vector) and the covariance function (covariance matrix) of the finite dimensional distribution of the induced Gaussian Process by the infinite-width tensor network with a general set-up. We study the properties of the covariance function and derive the approximation of the covariance function when the integral in the expectation operator is intractable. In the numerical experiments, we implement the Gaussian Process corresponding to the infinite limit tensor networks and plot the sample paths of these models. We study the hyperparameters and plot the sample path families in the induced Gaussian Process by varying the standard deviations of the prior distributions. As expected, the parameters in the prior distribution namely the hyper-parameters in the induced Gaussian Process controls the characteristic lengthscales of the Gaussian Process.
研究动机与目标
- 研究张量网络在宽度趋于无穷时的函数极限。
- 建立无限宽张量网络与高斯过程(GPs)之间的等价性。
- 在一般设置下,分析所得到GP的均值函数与协方差函数。
- 探讨先验分布的超参数(如标准差)如何影响所诱导GP的特征长度尺度。
- 通过在不同先验方差下对样本路径族进行数值实验,验证理论结果。
提出的方法
- 推导纯矩阵乘积态(MPS)的无限宽极限,证明其收敛于具有明确定义均值与协方差函数的GP。
- 将分析扩展至两种混合架构:神经核MPS与含隐藏神经层的MPS,证明其在无限宽极限下也收敛于GP。
- 使用期望算子计算所诱导GP的有限维分布的均值向量与协方差矩阵。
- 通过推导解析近似,处理协方差函数中难以计算的积分。
- 对无限宽张量网络所诱导的GP进行数值实现,并生成样本路径以供可视化。
- 系统性地改变先验分布的标准差,研究其对样本路径复杂度与长度尺度的影响。
实验结果
研究问题
- RQ1纯矩阵乘积态(MPS)张量网络的无限宽极限是否收敛于高斯过程?
- RQ2混合张量网络架构——特别是神经核MPS与含隐藏神经层的MPS——在无限宽极限下的行为如何?
- RQ3由无限宽张量网络所诱导的GP的均值函数与协方差函数的显式形式是什么?
- RQ4如先验分布的标准差等超参数如何影响所得到GP的特征长度尺度?
- RQ5数值实验能否证实增加先验方差会导致所诱导GP中样本路径更复杂且更具灵活性?
主要发现
- 纯矩阵乘积态(MPS)的无限宽极限收敛于高斯过程,但该GP是平凡的,因缺乏非线性而具有零不确定带。
- 神经核MPS与含隐藏神经层的MPS在宽度趋于无穷时均收敛于高斯过程,将GP等价性的适用范围从纯线性模型扩展至更广泛的情形。
- 在一般设置下,所诱导GP的均值函数与协方差函数被显式推导出来,为极限过程提供了闭式表征。
- 当期望积分难以计算时,对协方差函数进行近似,从而实现GP核的实际计算。
- 数值实验证实,增加先验分布的标准差会导致样本路径更复杂且变异性更高,直接控制GP的特征长度尺度。
- 从无限宽张量网络模型生成的样本路径族表明,先验方差等超参数控制函数空间的复杂度与平滑性,验证了对GP特性的理论调控能力。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。