Skip to main content
QUICK REVIEW

[论文解读] Adaptive variational Bayes: Optimality, computation and applications

Ilsang Ohn, Lizhen Lin|arXiv (Cornell University)|Sep 7, 2021
Model Reduction and Neural Networks参考文献 107被引用 4
一句话总结

本文提出了一种自适应变分贝叶斯框架,通过使用最优权重将多个模型的变分后验进行组合,从而在一般条件下实现自适应的后验收缩率。该方法在包含稀疏模型和深度学习模型的大规模模型集合中,既保证了计算上的可行性,又实现了最优的频率学派性能。

ABSTRACT

In this paper, we explore adaptive inference based on variational Bayes. Although several studies have been conducted to analyze the contraction properties of variational posteriors, there is still a lack of a general and computationally tractable variational Bayes method that performs adaptive inference. To fill this gap, we propose a novel adaptive variational Bayes framework, which can operate on a collection of models. The proposed framework first computes a variational posterior over each individual model separately and then combines them with certain weights to produce a variational posterior over the entire model. It turns out that this combined variational posterior is the closest member to the posterior over the entire model in a predefined family of approximating distributions. We show that the adaptive variational Bayes attains optimal contraction rates adaptively under very general conditions. We also provide a methodology to maintain the tractability and adaptive optimality of the adaptive variational Bayes even in the presence of an enormous number of individual models, such as sparse models. We apply the general results to several examples, including deep learning and sparse factor models, and derive new and adaptive inference results. In addition, we characterize an implicit regularization effect of variational Bayes and show that the adaptive variational posterior can utilize this.

研究动机与目标

  • 解决缺乏一种通用且计算上可行的变分贝叶斯方法,以在多个模型间实现自适应推断的问题。
  • 开发一种框架,无需依赖模型特定的先验分布或变分族,即可保持最优的后验收缩率。
  • 在处理大量模型(如稀疏或高维设置)时,确保计算上的可行性。
  • 在深度神经网络和稀疏因子模型等复杂模型中,展示该方法的自适应性与最优性。
  • 刻画变分贝叶斯的隐式正则化效应,并说明其如何在自适应框架中被利用。

提出的方法

  • 对模型集合中的每个独立模型,使用标准变分推断计算其各自的变分后验。
  • 通过最优权重将这些独立的变分后置分布组合,形成全局的自适应变分后验。
  • 所得的自适应变分后验是真实后验在预定义分布族中的最近似近似。
  • 采用两阶段流程:首先对每个模型进行变分推断,然后通过加权平均实现模型组合。
  • 借鉴Zhang & Gao (2020) 的理论结果,在较弱正则性条件下确保收缩率的最优性。
  • 利用变分贝叶斯的隐式正则化效应,提升高维和稀疏模型中的性能。

实验结果

研究问题

  • RQ1是否存在一种通用且计算上可行的变分贝叶斯方法,可在多种模型间实现自适应的后验收缩率?
  • RQ2如何最优地组合多个模型的变分后置分布,以逼近真实后验?
  • RQ3当模型数量庞大或处于高维设置时,所提出的框架是否仍能保持自适应最优性?
  • RQ4变分贝叶斯中的隐式正则化起什么作用,如何在自适应框架中加以利用?
  • RQ5该方法是否能在深度神经网络和稀疏因子模型等复杂模型中实现最优收缩率?

主要发现

  • 自适应变分后验在非常一般条件下实现了最优的后验收缩率,包括在非参数和高维设置中。
  • 即使在模型数量庞大的情况下(如稀疏模型集合),该方法仍保持计算上的可行性。
  • 与基于模型选择的变分贝叶斯相比,该框架能更紧密地逼近真实后验。
  • 自适应变分后验继承并利用了变分贝叶斯的隐式正则化效应,提升了估计的稳定性。
  • 理论结果在深度学习和稀疏因子模型中得到验证,获得了具有最优率的新自适应推断结果。
  • 该方法在频率学派风险下被证明是最优的,其收缩率在弱假设下与极小化最大风险的最优率一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。