Skip to main content
QUICK REVIEW

[论文解读] Measuring support for a hypothesis about a random parameter without estimating its unknown prior

David R. Bickel|arXiv (Cornell University)|Dec 31, 2010
Gene expression and cancer classification参考文献 43被引用 8
一句话总结

该论文提出了一种方法,用于在不需估计未知先验分布的情况下,衡量关于随机参数假设的统计支持。通过使用信息论导出的极小化极大最优度量——特别是基于归一化最大似然的方法——该方法在估计后验与先验对数似然比差异时实现了渐近无偏性,即使在数据有限的情况下(如蛋白质组学中仅有一个蛋白质)也表现出色。

ABSTRACT

For frequentist settings in which parameter randomness represents variability rather than uncertainty, the ideal measure of the support for one hypothesis over another is the difference in the posterior and prior log odds. For situations in which the prior distribution cannot be accurately estimated, that ideal support may be replaced by another measure of support, which may be any predictor of the ideal support that, on a per-observation basis, is asymptotically unbiased. Two qualifying measures of support are defined. The first is minimax optimal with respect to the population and is equivalent to a particular Bayes factor. The second is worst-sample minimax optimal and is equivalent to the normalized maximum likelihood. It has been extended by likelihood weights for compatibility with more general models. One such model is that of two independent normal samples, the standard setting for gene expression microarray data analysis. Applying that model to proteomics data indicates that support computed from data for a single protein can closely approximate the estimated difference in posterior and prior odds that would be available with the data for 20 proteins. This suggests the applicability of random-parameter models to other situations in which the parameter distribution cannot be reliably estimated.

研究动机与目标

  • 开发一种关于随机参数假设的统计支持度量,其无需估计未知先验分布。
  • 解决在样本量小或存在依赖关系的高维生物数据(如基因表达或蛋白质组学)中可靠估计先验分布的挑战。
  • 提供一种支持度量,其渐近地模拟后验与先验对数似然比的差异,确保可解释性并兼容贝叶斯推理。
  • 提供一种客观的、类似频率学派的证据度量,保留贝叶斯因子的可解释性,同时避免对主观或错误指定先验的依赖。
  • 使用真实蛋白质组学数据,将所提出的支撑度量与经验贝叶斯基准进行对比评估,特别是在仅有一个蛋白质数据可用时。

提出的方法

  • 将支持定义为后验与先验对数似然比的差异,这是一种标准的贝叶斯度量,但其公式化不依赖于先验分布的知识。
  • 引入两种极小化极大最优的支持度量:一种针对总体是最小化极大的,另一种针对最坏样本是最小化极大的。
  • 第一种度量等价于特定的贝叶斯因子,而第二种度量等价于归一化最大似然(NML),并通过似然加权推广至一般模型。
  • 将基于NML的度量应用于两个独立正态样本模型,这是微阵列和蛋白质组学数据分析中的标准设置。
  • 利用信息论原理(如最小描述长度)推导出渐近无偏的每观测度量,以支持度量的构建。
  • 采用涉及预测密度积分的惩罚因子,以防止在高维设置下过拟合。

实验结果

研究问题

  • RQ1能否定义一种不需估计随机参数未知先验分布的统计支持度量?
  • RQ2如何量化假设的支持度,使其在渐近意义上近似后验与先验对数似然比的差异,且无需先验知识?
  • RQ3此类支持度量在极小化极大准则下的最优性特征为何?其与最大似然比或经验贝叶斯等现有方法相比表现如何?
  • RQ4从单个蛋白质数据计算的支持度,能在多大程度上近似使用20个蛋白质数据可获得的后验-先验对数似然比差异?
  • RQ5在高维或小样本设置下,所提出的极小化极大支持度量相对于上界(最大似然比)的表现如何?

主要发现

  • 基于归一化最大似然(NML)的支持度量在最坏样本意义下是最小化极大的,且在先验分布未知时仍表现出稳健性。
  • 对于单个蛋白质,所提出的支持度量能非常接近使用20个蛋白质数据估计的后验-先验对数似然比差异。
  • 与作为上界的最大似然比相比,基于NML的支持度量更能有效避免过拟合。
  • 极小化极大支持度量保持保守性,不会像上界那样在高维设置下过度估计后验与先验对数似然比的真实差异。
  • 该方法适用于经验贝叶斯先验无法可靠估计的模型,例如样本有限或特征相关性高的蛋白质组学场景。
  • 该支持度量在不同样本量下均具有可解释性,无需校准,因为它直接量化了从先验到后验的几率变化,与贝叶斯解释一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。