Skip to main content
QUICK REVIEW

[论文解读] Large-scale interval and point estimates from an empirical Bayes extension of confidence posteriors

David R. Bickel|arXiv (Cornell University)|Dec 29, 2010
Gene expression and cancer classification参考文献 30被引用 10
一句话总结

本文提出了一种置信后验的半参数经验贝叶斯扩展,将局部错误发现率(LFDR)估计与置信分布相结合,用于高维生物数据中的收缩区间和点估计。通过结合基于LFDR的收缩与频率学派的覆盖性质,该方法生成了更短、更准确的置信区间,其覆盖真实参数的概率高于名义水平,从而在基因组学和生物信息学中提升了特征优先排序的效果。

ABSTRACT

The proposed approach extends the confidence posterior distribution to the semi-parametric empirical Bayes setting. Whereas the Bayesian posterior is defined in terms of a prior distribution conditional on the observed data, the confidence posterior is defined such that the probability that the parameter value lies in any fixed subset of parameter space, given the observed data, is equal to the coverage rate of the corresponding confidence interval. A confidence posterior that has correct frequentist coverage at each fixed parameter value is combined with the estimated local false discovery rate to yield a parameter distribution from which interval and point estimates are derived within the framework of minimizing expected loss. The point estimates exhibit suitable shrinkage toward the null hypothesis value, making them practical for automatically ranking features in order of priority. The corresponding confidence intervals are also shrunken and tend to be much shorter than their fixed-parameter counterparts, as illustrated with gene expression data. Further, simulations confirm a theoretical argument that the shrunken confidence intervals cover the parameter at a higher-than-nominal frequency.

研究动机与目标

  • 为解决高维推断中缺乏一致的区间估计问题,特别是错误发现率(FDR)和局部错误发现率(LFDR)估计的问题。
  • 开发一种方法,生成与LFDR估计背后的层次模型一致的收缩点估计和区间估计,且无需强参数假设。
  • 通过使用基于置信后验的后验中位数对特征进行排序,改进基因组学中的特征优先排序。
  • 通过利用置信后验框架,确保即使在更短的置信区间下,频率学派的覆盖性质仍能被保持或增强。
  • 为传统置信区间和点估计提供一种实用的替代方案,这些方法在高维设置下无法反映不确定性和收缩效应。

提出的方法

  • 该方法将置信后验——由嵌套置信区间的覆盖概率定义——扩展至经验贝叶斯框架,使用估计的LFDR作为权重。
  • 为每个参数 $ \theta_i $ 构建边际置信后验 $ \hat{P}^{(i)} $,结合置信分布与估计的LFDR $ \hat{\ell}_i $。
  • 使用 $ \hat{P}^{(i)} $ 的后验中位数作为收缩点估计,该估计具有参数化不变性,并可调整统计显著性。
  • 置信区间以该后验中位数为中心,向原假设收缩,相比固定参数区间显著减小宽度。
  • 该方法确保收缩后区间具有高于名义水平的覆盖率,经理论推导和模拟验证。
  • 该方法通过在半参数经验贝叶斯设定下使用LFDR估计作为先验信息的代理,避免了对完整先验分布的依赖。

实验结果

研究问题

  • RQ1能否将置信后验扩展至经验贝叶斯框架,以在高维生物数据中生成一致的、基于收缩的区间估计与点估计?
  • RQ2尽管区间宽度减小,这些结果置信区间是否仍能保持或超过名义覆盖率?
  • RQ3扩展后的置信后验的后验中位数能否作为可靠且参数化不变的排序标准,用于优先排序生物特征?
  • RQ4与传统置信区间和点估计相比,该方法在估计准确性和区间宽度方面的表现如何?
  • RQ5在存在相关性和高维性的情况下,整合LFDR估计在多大程度上提升了区间估计的可靠性?

主要发现

  • 模拟结果证实,由所提置信后验导出的收缩置信区间覆盖真实参数值的概率高于名义的95%水平。
  • 置信后验的后验中位数提供了稳健且参数化不变的点估计,在命中或紧密逼近真实参数值方面表现良好。
  • 基于边际置信后验 $ \hat{P}^{(i)} $ 的置信区间显著短于固定参数区间,同时保持或提升了覆盖性能。
  • 该方法成功地将LFDR估计整合进频率学派框架,而无需对参数分布作强参数假设。
  • 在基因表达数据中,该方法生成了更短、更具信息量的置信区间,且以收缩估计为中心,从而增强了特征优先排序效果。
  • 该方法避免了将LFDR解释为后验概率的陷阱,认识到LFDR估计是保守的,不应被视为确定性结果。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。