[论文解读] Precise characterization of the prior predictive distribution of deep ReLU networks
本文利用梅杰-G函数,对有限宽度深度ReLU网络的先验预测分布进行了精确的解析表征,揭示了深度如何增强重尾性,而宽度则促进高斯性。该研究建立了与无限深度极限的正式联系,表明其收敛于正态-对数正态混合分布,并提出了基于深度和宽度的广义He先验,可直接控制预测方差。
Recent works on Bayesian neural networks (BNNs) have highlighted the need to better understand the implications of using Gaussian priors in combination with the compositional structure of the network architecture. Similar in spirit to the kind of analysis that has been developed to devise better initialization schemes for neural networks (cf. He- or Xavier initialization), we derive a precise characterization of the prior predictive distribution of finite-width ReLU networks with Gaussian weights. While theoretical results have been obtained for their heavy-tailedness, the full characterization of the prior predictive distribution (i.e. its density, CDF and moments), remained unknown prior to this work. Our analysis, based on the Meijer-G function, allows us to quantify the influence of architectural choices such as the width or depth of the network on the resulting shape of the prior predictive distribution. We also formally connect our results to previous work in the infinite width setting, demonstrating that the moments of the distribution converge to those of a normal log-normal mixture in the infinite depth limit. Finally, our results provide valuable guidance on prior design: for instance, controlling the predictive variance with depth- and width-informed priors on the weights of the network.
研究动机与目标
- 理解有限宽度深度ReLU网络中先验预测分布的分布特性,尽管已知其具有重尾行为,但其特性仍缺乏精确刻画。
- 形式化架构超参数(尤其是深度和宽度)对先验预测分布形状的影响。
- 将无限宽度极限分析扩展至无限深度情形,恢复并推广了关于神经网络高斯过程的先前结果。
- 提供一个原理性框架,用于设计可解释的、函数空间中方差感知的先验,如广义He先验。
提出的方法
- 作者采用梅杰-G函数,对任意深度和有限宽度的深度ReLU网络中先验预测分布的概率密度函数进行解析表征。
- 利用梅林变换和涉及Gamma函数与二项式系数的组合恒等式,推导出先验预测分布的矩的精确表达式。
- 该分析利用了带有ReLU非线性的随机矩阵乘积结构,将问题简化为具有随机权重的中间层激活之和。
- 通过分析矩生成函数的渐近行为,研究了无限深度极限,表明其收敛于正态-对数正态混合分布。
- 该方法可推导出广义He先验,通过指定与深度和宽度相关的权重方差,直接控制预测方差。
实验结果
研究问题
- RQ1有限宽度深度ReLU网络的先验预测分布如何依赖于其深度和宽度?
- RQ2此类网络中先验预测密度、累积分布函数和矩的精确解析形式是什么?
- RQ3在无限深度极限下,先验预测分布的矩如何表现?其收敛于何种分布?
- RQ4我们能否设计出可直接控制函数空间中预测方差的权重先验,而无需对各层进行单独调优?
主要发现
- 通过梅杰-G函数,精确表征了深度ReLU网络的先验预测分布,实现了其密度、CDF和矩的完整解析描述。
- 网络越深,其分布越呈现重尾特性;而越宽,其行为越趋近高斯分布,且深度对尾部行为的影响更强。
- 在无限深度极限下,先验预测分布的矩收敛于正态-对数正态混合分布的矩,即使在非渐近区域也与经验观察一致。
- 当深度随宽度线性增长(l-1 = γm)时,偶数矩收敛于 (2k−1)!!·exp(5γk(k−1)/2),证实了对数正态成分的出现。
- 本文提出了广义He先验,可直接指定预测方差,相较于逐层方差调优,提供了更具可解释性的替代方案。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。