Skip to main content
QUICK REVIEW

[论文解读] A Corrected and More Efficient Suite of MCMC Samplers for the Multinomal Probit Model

Xiyun Jiao, David A. van Dyk|arXiv (Cornell University)|Apr 29, 2015
Markov Chains and Monte Carlo Methods参考文献 5被引用 4
一句话总结

本文纠正了广泛使用的多项式 probit 模型 MCMC 采样器中的关键错误,特别是 Imai 和 van Dyk(2005a)以及 Burgette 和 Nordheim(2012)提出的算法,这些错误会改变平稳分布并损害后验推断。作者通过适当的重新参数化与约束施加,提出了修正后的采样器,并通过模拟和真实数据验证,表明尽管计算成本略有增加,修正后的算法在后验估计准确性与混合效率方面均有显著提升。

ABSTRACT

The multinomial probit (MNP) model is a useful tool for describing discrete-choice data and there are a variety of methods for fitting the model. Among them, the algorithms provided by Imai and van Dyk (2005a), based on Marginal Data Augmentation, are widely used, because they are efficient in terms of convergence and allow the possibly improper prior distribution to be specified directly on identifiable parameters. Burgette and Nordheim (2012) modify a model and algorithm of Imai and van Dyk (2005a) to avoid an arbitrary choice that is often made to establish identifiability. There is an error in the algorithms of Imai and van Dyk (2005a), however, which affects both their algorithms and that of Burgette and Nordheim (2012). This error can alter the stationary distribution and the resulting fitted parameters as well as the efficiency of these algorithms. We propose a correction and use both a simulation study and a real-data analysis to illustrate the difference between the original and corrected algorithms, both in terms of their estimated posterior distributions and their convergence properties. In some cases, the effect on the stationary distribution can be substantial.

研究动机与目标

  • 识别并修正 Imai 和 van Dyk(2005a)MCMC 算法中的两个关键错误,这些错误影响平稳分布与后验推断。
  • 将修正扩展至 Burgette 和 Nordheim(2012)的改进算法,该算法继承了相同的错误。
  • 展示这些错误对后验估计、收敛性与有效样本量(ESS)的实际影响。
  • 提出修正后的 MCMC 采样器,在保持原始算法效率的同时确保正确的后验抽样。
  • 更新流行的 R 包 'MNP',集成修正后的算法,以防止错误推断的广泛使用。

提出的方法

  • 通过适当的重新参数化重新表达方差-协方差矩阵的抽样步骤,避免参数变换错误。
  • 在吉布斯抽样过程中强制施加方差-协方差矩阵为正定矩阵的约束,纠正原始算法中的疏漏。
  • 引入拒绝采样步骤,确保修正后的采样器能够正确地目标于后验分布。
  • 利用边际数据增强(MDA)框架,在保持计算效率的同时确保后验模拟的有效性。
  • 通过每秒有效样本量(ESS)与分位数-分位数图比较修正与原始算法,评估混合性与准确性。
  • 通过模拟研究与对 507 笔人造黄油购买数据的真实数据分析,验证修正后算法的有效性。

实验结果

研究问题

  • RQ1Imai 和 van Dyk(2005a)MCMC 采样器中的错误如何影响多项式 probit 模型的平稳分布与后验推断?
  • RQ2Burgette 和 Nordheim(2012)算法中的错误(该算法基于 Imai 和 van Dyk)在多大程度上损害了后验估计的有效性?
  • RQ3与原始版本相比,修正后的算法在收敛性与有效样本量(ESS)方面有何改善?
  • RQ4修正的计算成本是多少?更高的混合效率是否足以证明额外计算的合理性?
  • RQ5与原始缺陷算法相比,修正后的采样器在真实世界数据上的表现如何?

主要发现

  • Imai 和 van Dyk(2005a)以及 Burgette 和 Nordheim(2012)算法中的错误显著改变了平稳分布,导致后验估计偏差与标准误错误。
  • 算法 1.1(原始版本)无法生成目标后验分布的抽样结果,如分位数-分位数图所示;而修正后的算法 1.3 在自相关性与收敛性方面表现显著更优。
  • 尽管计算时间增加了约 15%,修正后的算法 1.3 每秒的 ESS 高于原始算法 1.1,表明其混合效率更高。
  • 修正后的算法 2.2 和 3.2 与原始版本(2.1 和 3.1)性能相似,但仅修正版本能生成有效的后验样本。
  • 对 507 笔人造黄油购买数据的真实数据分析表明,原始算法产生错误的后验分布,而修正版本则生成可靠且混合良好的马尔可夫链。
  • 作者确认,修正后的算法仅需约 15% 的额外计算时间,但能实现显著更高的估计准确性与效率。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。