[论文解读] Richer priors for infinitely wide multi-layer perceptrons
该论文通过引入更丰富的先验分布——非零均值和部分可交换性(RCE)先验,扩展了无限宽多层感知机(MLP)与高斯过程(GP)之间经典的联系。研究表明,此类先验可导出具有超参数层次推断的GP极限,从而产生更具表达力的核函数,避免深层网络中的病态问题,尽管边际似然函数随深度增加而变得日益病态。
It is well-known that the distribution over functions induced through a zero-mean iid prior distribution over the parameters of a multi-layer perceptron (MLP) converges to a Gaussian process (GP), under mild conditions. We extend this result firstly to independent priors with general zero or non-zero means, and secondly to a family of partially exchangeable priors which generalise iid priors. We discuss how the second prior arises naturally when considering an equivalence class of functions in an MLP and through training processes such as stochastic gradient descent. The model resulting from partially exchangeable priors is a GP, with an additional level of inference in the sense that the prior and posterior predictive distributions require marginalisation over hyperparameters. We derive the kernels of the limiting GP in deep MLPs, and show empirically that these kernels avoid certain pathologies present in previously studied priors. We empirically evaluate our claims of convergence by measuring the maximum mean discrepancy between finite width models and limiting models. We compare the performance of our new limiting model to some previously discussed models on synthetic regression problems. We observe increasing ill-conditioning of the marginal likelihood and hyper-posterior as the depth of the model increases, drawing parallels with finite width networks which require notoriously involved optimisation tricks.
研究动机与目标
- 将经典结论推广:无限宽MLP在i.i.d.零均值权重下收敛于GP,通过允许非零均值先验。
- 引入一种部分可交换先验(RCE),其源于MLP中的置换对称性及SGD训练动力学。
- 推导在这些更丰富先验下深层MLP的极限GP核函数,并分析其性质。
- 通过合成回归任务的实验评估有限宽模型向极限GP的收敛性及性能表现。
- 研究在新先验结构下,深层模型中边际似然与超后验的病态性表现。
提出的方法
- 利用中心极限定理与泛函中心极限定理,推导在一般零均值或非零均值i.i.d.先验下深层MLP的极限GP核函数。
- 引入一种部分可交换先验(RCE),通过允许层内依赖关系而推广i.i.d.先验,同时保持对称性。
- 证明所得模型为具有层次先验的GP,需对超参数进行边际化处理(即二级推断)。
- 推导在ReLU与Leaky ReLU激活函数下,具有非零均值时的显式核函数表达式。
- 采用最大均值差异(MMD)度量有限宽模型向极限GP的收敛程度。
- 使用具有二元正态提议分布的Metropolis-Hastings采样器探索超后验与边际似然,并在均值与方差上采用非信息性超先验。
实验结果
研究问题
- RQ1在先验中允许非零均值,对深层MLP的极限GP核函数有何影响?
- RQ2能否从MLP的对称性与SGD训练动力学中推导出部分可交换先验(RCE)?其如何推广i.i.d.先验?
- RQ3在非零均值与RCE先验下,极限GP核函数的结构与功能特性为何?与标准i.i.d.先验相比有何差异?
- RQ4在新先验结构下,深层模型中的边际似然与超后验行为如何?随着深度增加是否会出现病态性?
- RQ5与标准GP先验相比,新核结构是否能在合成回归任务中提升泛化性能?
主要发现
- 在非零均值先验下,极限GP表现出更具表达力的核结构,尤其在深层网络中优于标准零均值情形。
- RCE先验自然源于MLP中的置换对称性及SGD训练动力学。
- 所得模型为具有层次推断的GP,其先验与后验预测分布需对超参数进行边际化处理。
- 实证评估显示,有限宽模型与极限GP之间的MMD随宽度增加而减小,证实了收敛性。
- 边际似然与超后验随深度增加而变得日益病态,与有限宽网络优化中的挑战一致。
- 具有非零均值的模型在深层架构中表现出更高的过拟合敏感性,表明表达力与泛化能力之间存在权衡。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。