Skip to main content
QUICK REVIEW

[论文解读] Being an informed Bayesian: Assessing prior informativeness and prior likelihood conflict

Matthew Reimherr, Xiao‐Li Meng|arXiv (Cornell University)|Jun 23, 2014
Statistical Methods and Inference被引用 11
一句话总结

本文提出一种诊断方法,用于评估贝叶斯分析中先验信息的丰富程度,并检测先验与似然之间的冲突。该方法通过量化为达到相同后验不确定性所需、给定先验与基线先验之间的有效样本量差异(M(k))来实现。研究发现,当存在冲突时,M(k) 随数据量增加而减少;并识别出一种超信息现象,即当先验与真实情况一致时,其获得额外的有效数据量。

ABSTRACT

Dramatically expanded routine adoption of the Bayesian approach has substantially increased the need to assess both the confirmatory and contradictory information in our prior distribution with regard to the information provided by our likelihood function. We propose a diagnostic approach that starts with the familiar posterior matching method. For a given likelihood model, we identify the difference in information needed to form two likelihood functions that, when combined respectively with a given prior and a baseline prior, will lead to the same posterior uncertainty. In cases with independent, identically distributed samples, sample size is the natural measure of information, and this difference can be viewed as the prior data size $M(k)$, with regard to a likelihood function based on $k$ observations. When there is no detectable prior-likelihood conflict relative to the baseline, $M(k)$ is roughly constant over $k$, a constant that captures the confirmatory information. Otherwise $M(k)$ tends to decrease with $k$ because the contradictory prior detracts information from the likelihood function. In the case of extreme contradiction, $M(k)/k$ will approach its lower bound $-1$, representing a complete cancelation of prior and likelihood information due to conflict. We also report an intriguing super-informative phenomenon where the prior effectively gains an extra $(1+r)^{-1}$ percent of prior data size relative to its nominal size when the prior mean coincides with the truth, where $r$ is the percentage of the nominal prior data size relative to the total data size underlying the posterior. We demonstrate our method via several examples, including an application exploring the effect of immunoglobulin levels on lupus nephritis. We also provide a theoretical foundation of our method for virtually all likelihood-prior pairs that possess asymptotic conjugacy.

研究动机与目标

  • 为常规贝叶斯应用中日益增长的先验信息丰富程度系统评估需求提供解决方案。
  • 检测可能损害后验可靠性、先验分布与似然函数之间的冲突。
  • 量化为抵消先验-似然不一致对后验不确定性的影响,所需额外数据的量。
  • 探讨先验在何种条件下可能表现出悖论性地比其名义大小更具信息量。
  • 为具有渐近共轭性的广泛类别的似然-先验对提供理论基础。

提出的方法

  • 使用后验匹配方法,在相同似然下使给定先验与基线先验的后验不确定性相等。
  • 将 M(k) 定义为使后验方差相同,所需给定先验与基线先验之间有效样本量的差异。
  • 以独立同分布数据的样本量为信息度量标准,将 M(k) 视为观测样本量 k 的函数。
  • 分析 M(k) 随 k 的行为:若 M(k) 恒定,表示先验为确认性;若 M(k) 减小,则表明存在先验-似然冲突。
  • 推导出 M(k)/k 趋近于 -1 的理论条件,表示由于冲突导致信息完全抵消。
  • 识别出一种超信息现象:当先验均值与真实值一致时,先验的有效数据量会增加 (1+r)⁻¹ 百分比,其中 r 为名义先验大小与总数据量的比值。

实验结果

研究问题

  • RQ1我们如何正式评估一个先验是否相对于似然提供确认性或矛盾性信息?
  • RQ2给定先验与基线先验之间,导致相同后验不确定性的有效样本量差异 M(k) 是什么?
  • RQ3随着观测样本量 k 增加,M(k) 如何变化?这揭示了关于先验-似然相容性的何种信息?
  • RQ4在何种条件下,即使先验均值正确,先验仍可能显得比其名义大小更具信息量?
  • RQ5何种理论条件可确保该方法在广泛类别的似然-先验对中保持有效?

主要发现

  • 当不存在先验-似然冲突时,M(k) 在 k 的取值范围内保持近似恒定,表明先验确认了似然提供的信息。
  • 当存在冲突时,M(k) 随 k 增加而减少,反映出由于先验错位,似然信息的贡献逐渐减弱。
  • 在极端冲突情形下,M(k)/k 趋近于 -1,表示先验与似然信息完全相互抵消。
  • 观察到一种超信息现象:当先验均值与真实参数值一致时,先验的有效数据量会增加 (1+r)⁻¹ 百分比。
  • 该方法在几乎所有具有渐近共轭性的似然-先验对中均具有理论基础,确保了广泛适用性。
  • 该方法在真实世界应用中得到验证,涉及免疫球蛋白水平与狼疮性肾炎的研究,展示了其在生物医学研究中的实际效用。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。