Skip to main content
QUICK REVIEW

[论文解读] A framework for posterior consistency in model selection

David Rossell|arXiv (Cornell University)|Jun 11, 2018
Statistical Methods and Inference被引用 3
一句话总结

本文提出了一种有限样本框架,用于分析线性回归中的贝叶斯模型选择,将先验选择与问题特征——如样本量、信噪比、维度和稀疏性——联系起来,探讨后验一致性与模型选择准确性。该研究建立了 $L_0$ 惩罚和伪后验的有限样本速率,揭示了在稀疏先验和模型误设下,即使样本量中等,渐近最优性也可能因效能损失而产生误导,强调应基于数据特定特征而非仅依赖渐近行为来评估方法性能。

ABSTRACT

We develop a framework to help understand frequentist properties of Bayesian model selection, specifically its ability to select the (Kullback-Leibler) optimal model and portray model selection uncertainty. We outline its general basis and then focus on linear regression. The contribution is not proving consistency under given prior conditions but providing finite-sample rates that describe how model selection depends on the prior and problem characteristics such as sample size, signal-to-noise, problem dimension and true sparsity. A corollary proves a strong form of convergence for $L_0$ penalties and pseudo-posterior probabilities of interest for $L_0$ uncertainty quantification. These results unify and extend current Bayesian model selection literature and signal limitations, specifically that asymptotically optimal sparse priors can significantly reduce power even for moderately large $n$ and that less sparse priors can improve power trade-offs not adequately captured by asymptotic rates. These issues are compounded by the fact that model misspecification often causes an exponential drop in power, as we briefly study here. Our examples confirm these findings, underlining the importance of considering the data at hand's characteristics to judge the quality of model selection procedures, rather than relying purely on asymptotics.

研究动机与目标

  • 开发一种有限样本框架,以理解贝叶斯模型选择的频率性质,特别是后验一致性和不确定性量化。
  • 研究先验分布以及问题特定特征(如样本量、信噪比、维度和真实稀疏性)如何影响模型选择性能。
  • 通过证明渐近最优先验在有限样本中可能导致效能显著下降,挑战对渐近一致性结果的依赖。
  • 表明在中等样本中,较不稀疏的先验可能提供优于渐近速率预测的效能权衡。
  • 考察模型误设对效能的影响,揭示性能出现指数级下降。

提出的方法

  • 该框架通过推导依赖于关键问题特征的有限样本速率,分析贝叶斯模型选择中的后验一致性。
  • 聚焦于线性回归,研究 $L_0$ 惩罚和伪后验概率在不确定性量化中的应用。
  • 该方法建立了 $L_0$ 基后验的强收敛形式,将其行为与先验超参数及数据结构相联系。
  • 在有限样本条件下,比较不同先 priors 类别(特别是稀疏与较不稀疏先验)的性能。
  • 将信噪比、维度和真实稀疏性作为评估模型选择准确性的关键协变量纳入分析。
  • 通过理论推导和示例说明,展示渐近一致性在指导有限样本模型选择时的局限性。

实验结果

研究问题

  • RQ1先验分布以及样本量、稀疏性等问题特征如何影响贝叶斯模型选择的有限样本性能?
  • RQ2渐近最优的稀疏先验在中等样本中在多大程度上会损害效能,原因是什么?
  • RQ3较不稀疏的先验是否可能提供优于渐近速率预测的效能权衡,以及在何种条件下?
  • RQ4模型误设如何影响贝叶斯模型选择程序的效能?
  • RQ5在线性回归中,$L_0$ 惩罚后验和伪后验的收敛由哪些有限样本速率决定?

主要发现

  • 贝叶斯模型选择中后验一致性的有限样本速率,关键取决于样本量、信噪比、维度和真实稀疏性。
  • 即使在中等偏大的样本量下,渐近最优的稀疏先验仍可能导致显著的效能损失,削弱其实际应用价值。
  • 在有限样本中,较不稀疏的先验通常能提供优于渐近理论预测的效能权衡。
  • 模型误设导致效能出现指数级下降,而这一现象无法被渐近一致性结果充分捕捉。
  • 证明了 $L_0$ 惩罚和伪后验概率的强收敛形式,统一并扩展了现有的贝叶斯模型选择理论。
  • 研究表明,必须直接评估数据特定特征,而非仅依赖渐近性质来判断模型选择质量。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。