Skip to main content
QUICK REVIEW

[论文解读] Bayesian Cross Validation and WAIC for Predictive Prior Design in Regular Asymptotic Theory

Sumio Watanabe|arXiv (Cornell University)|Mar 27, 2015
Advanced Multi-Objective Optimization Algorithms参考文献 9被引用 4
一句话总结

本文为在常规统计模型中使用贝叶斯交叉验证(CV)和WAIC优化超先验提供了理论基础。证明了最小化CV或WAIC在渐近意义上可最小化平均泛化损失,提供了将这些准则表示为超参数函数的直接公式,从而实现高效且稳定的超参数优化,并降低方差。

ABSTRACT

Prior design is one of the most important problems in both statistics and machine learning. The cross validation (CV) and the widely applicable information criterion (WAIC) are predictive measures of the Bayesian estimation, however, it has been difficult to apply them to find the optimal prior because their mathematical properties in prior evaluation have been unknown and the region of the hyperparameters is too wide to be examined. In this paper, we derive a new formula by which the theoretical relation among CV, WAIC, and the generalization loss is clarified and the optimal hyperparameter can be directly found. By the formula, three facts are clarified about predictive prior design. Firstly, CV and WAIC have the same second order asymptotic expansion, hence they are asymptotically equivalent to each other as the optimizer of the hyperparameter. Secondly, the hyperparameter which minimizes CV or WAIC makes the average generalization loss to be minimized asymptotically but does not the random generalization loss. And lastly, by using the mathematical relation between priors, the variances of the optimized hyperparameters by CV and WAIC are made smaller with small computational costs. Also we show that the optimized hyperparameter by DIC or the marginal likelihood does not minimize the average or random generalization loss in general.

研究动机与目标

  • 解决在使用CV和WAIC进行超参数优化时存在的理论不一致性,特别是其与平均泛化损失关系的问题。
  • 应对由于需在广阔超参数空间内计算后验分布而导致的超参数搜索计算成本过高的实际挑战。
  • 阐明通过CV、WAIC、DIC和边际似然进行超参数优化的渐近行为,尤其关注其与最小化泛化误差的关系。
  • 提出一种直接估计CV和WAIC作为超参数函数的方法,从而实现无需全面比较候选解的高效优化。
  • 通过先验与自平均性质之间的数学关系,降低优化后超参数的方差。

提出的方法

  • 推导出一个新理论公式,将CV和WAIC明确表示为超参数的函数,从而可在无需重复后验采样的情况下直接计算。
  • 应用常规渐近理论分析CV和WAIC的二阶展开,证明其在最小化平均泛化损失方面具有等价性。
  • 利用奇异学习理论和发散参数的概念,识别出CV和WAIC可能无法取得最小值的条件。
  • 引入自平均方法,通过在多个数据集上估计CV和WAIC的期望值,降低优化后超参数的方差。
  • 通过理论渐近分析,比较CV、WAIC、DIC和边际似然在最小化泛化损失方面的性能。
  • 提出通过从预测分布或经验分布进行蒙特卡洛采样,对WAICRS(自平均WAIC)进行数值估计,以提高稳定性。

实验结果

研究问题

  • RQ1在常规模型中,CV和WAIC作为超参数优化器是否渐近等价?
  • RQ2最小化CV或WAIC是否渐近最小化平均泛化损失?这与最小化随机泛化损失(依赖于特定训练集)有何区别?
  • RQ3能否通过新解析公式直接找到最优超参数,而无需全面搜索?
  • RQ4通过CV和WAIC优化的超参数方差如何比较?能否通过自平均方法降低其方差?
  • RQ5为何DIC和边际似然即使在渐近意义上也无法最小化平均泛化损失?

主要发现

  • CV和WAIC具有相同的二阶渐近展开,因此在常规模型中作为超参数优化器具有渐近等价性。
  • 最小化CV或WAIC可渐近最小化平均泛化损失,但无法最小化随机泛化损失(后者依赖于特定训练集)。
  • 通过CV或WAIC找到的最优超参数收敛于最小化期望泛化误差的参数,而非单个数据集上观测到的误差。
  • 新解析公式允许直接计算CV和WAIC作为超参数的函数,从而无需全面候选评估,显著降低计算成本。
  • 通过应用自平均技术,可降低优化后超参数的方差,尤其WAICRS表现出低于标准CV或WAIC的方差。
  • 通过DIC或边际似然优化的超参数无法渐近最小化平均泛化损失,而CV或WAIC则可以。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。