Skip to main content
QUICK REVIEW

[论文解读] Block Hyper-g Priors in Bayesian Regression

Agniva Som, Christopher M. Hans|arXiv (Cornell University)|Jun 25, 2014
Bayesian Methods and Mixture Models参考文献 49被引用 6
一句话总结

本文提出了一种新型贝叶斯回归先验——块超g先验(block hyper-g prior),该先验将预测变量划分为多个块,并对每个块独立应用超g先验,从而避免了传统尺度混合g先验中出现的不良行为,如本质上最小二乘法(Essentially Least Squares, ELS)估计和条件林德利悖论(Conditional Lindley’s Paradox, CLP)。该方法在块正交设计和渐近情形下可保证模型选择与预测的一致性。

ABSTRACT

The development of prior distributions for Bayesian regression has traditionally been driven by the goal of achieving sensible model selection and parameter estimation. The formalization of properties that characterize good performance has led to the development and popularization of thick tailed mixtures of g priors such as the Zellner--Siow and hyper-g priors. The properties of a particular prior are typically illuminated under limits on the likelihood or the prior. In this paper we introduce a new, conditional information asymptotic that is motivated by the common data analysis setting where at least one regression coefficient is much larger than others. We analyze existing mixtures of g priors under this limit and reveal two new behaviors, Essentially Least Squares (ELS) estimation and the Conditional Lindley's Paradox (CLP), and argue that these behaviors are, in general, undesirable. As the driver behind both of these behaviors is the use of a single, latent scale parameter that is common to all coefficients, we propose a block hyper-g prior, defined by first partitioning the covariates into groups and then placing independent hyper-g priors on the corresponding blocks of coefficients. We provide conditions under which ELS and the CLP are avoided by the new class of priors, and provide consistency results under traditional sample size asymptotics.

研究动机与目标

  • 为解决现有尺度混合g先验在某一回归系数远大于其他系数时的局限性。
  • 识别并形式化传统g先验框架中的两种不良行为——本质上最小二乘法(ELS)估计与条件林德利悖论(CLP)。
  • 提出一类新先验——块超g先验,通过为每个预测变量块分配独立的尺度参数,避免ELS与CLP。
  • 在条件信息渐近和传统样本量渐近两种情形下,建立块超g先验的理论一致性。
  • 证明该先验在保持计算效率的同时,可在块正交设计下提升模型选择与预测性能。

提出的方法

  • 通过将协变量划分为互不相交的块,并为每个块的回归系数分配独立的超g先验,构建块超g先验。
  • 每个块的系数被赋予一个均值为零、方差由独立g_i缩放的正态先验,并在g_i参数上设置超先验以实现收缩。
  • 利用设计矩阵中的块正交性,确保解析可处理性,并推导后验矩与模型概率。
  • 在条件信息渐近框架下进行理论分析,即当某一系数相对于其他系数趋于无穷大时,研究其极限行为。
  • 在模型平均与边际似然方法下,推导出模型选择与预测的一致性结果。
  • 通过块正交设计设置验证理论性质,实现后验期望与模型包含概率的显式计算。

实验结果

研究问题

  • RQ1当某一回归系数显著大于其他系数时,标准超g先验是否表现出不良的极限行为?
  • RQ2本质上最小二乘法(ELS)估计与条件林德利悖论(CLP)在多大程度上破坏了现有g先验框架下的模型选择与估计?
  • RQ3是否可通过引入多个、块特定的尺度参数的先验结构,消除高杠杆系数情形下的ELS与CLP现象?
  • RQ4在何种条件下,块超g先验能在大系数存在时实现模型选择与预测的一致性?
  • RQ5在条件信息极限下,块超g先延与标准g先验在渐近行为与性能表现上如何比较?

主要发现

  • 标准超g先验及其他尺度混合g先验在条件信息渐近下表现出本质上最小二乘法(ELS)估计,即后验估计坍缩为最小二乘估计。
  • 这些先验还遭受条件林德利悖论(CLP)的影响,导致当真实系数远大于其他系数时,模型选择不一致。
  • 块超g先延通过为每个块独立使用g先验,避免了ELS与CLP,从而防止单一全局尺度参数的主导作用。
  • 在块正交设计下,块超g先验实现了模型选择一致性,后验模型概率随样本量增加而集中于真实模型。
  • 在贝叶斯模型平均(BMA)下,该先验确保了预测一致性,模型平均预测器在极限下收敛于真实条件均值。
  • 理论结果证实,块超g先验在保持理想收缩与估计性质的同时,避免了单尺度g先延的病态行为。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。