[论文解读] Block-Wise Pseudo-Marginal Metropolis-Hastings
本文提出了一种分块伪边缘后验马尔可夫链蒙特卡洛方法,通过将随机数划分为多个区块,使当前参数与提议参数的似然估计仅在一个区块上存在差异,从而在贝叶斯推断中提升效率。该方法增强了对数似然估计之间的相关性,简化了理论分析,并在面板数据和子采样应用中实现了显著的速度提升,且在最优粒子数选择下效果更佳。
The pseudo-marginal Metropolis-Hastings approach is increasingly used for Bayesian inference in statistical models where the likelihood is analytically intractable but can be estimated unbiasedly, such as random effects models and state-space models, or for data subsampling in big data settings. In a seminal paper, Deligiannidis et al. (2015) show how the pseudo-marginal Metropolis-Hastings (PMMH) approach can be made much more e cient by correlating the underlying random numbers used to form the estimate of the likelihood at the current and proposed values of the unknown parameters. Their proposed approach greatly speeds up the standard PMMH algorithm, as it requires a much smaller number of particles to form the optimal likelihood estimate. We present a closely related alternative PMMH approach that divides the underlying random numbers mentioned above into blocks so that the likelihood estimates for the proposed and current values of the likelihood only di er by the random numbers in one block. Our approach is less general than that of Deligiannidis et al. (2015), but has the following advantages. First, it provides a more direct way to control the correlation between the logarithms of the estimates of the likelihood at the current and proposed values of the parameters. Second, the mathematical properties of the method are simplified and made more transparent compared to the treatment in Deligiannidis et al. (2015). Third, blocking is shown to be a natural way to carry out PMMH in, for example, panel data models and subsampling problems. We obtain theory and guidelines for selecting the optimal number of particles, and document large speed-ups in a panel data example and a subsampling problem.
研究动机与目标
- 开发一种比现有不可行似然伪边缘MCMC方法更透明且高效的替代方法。
- 通过系统性地对随机数进行分块,提高当前与提议参数值下对数似然估计之间的相关性。
- 与先前采用相关随机数的方法相比,简化伪边缘算法的理论分析。
- 为分块伪边缘后验马尔可夫链蒙特卡洛方法提供最优粒子数选择的实用指导。
- 在面板数据模型和大规模数据子采样问题中展示显著的计算加速效果。
提出的方法
- 该方法将似然估计中使用的独立同分布随机数集合划分为互不相交的区块。
- 在每次MCMC迭代中,仅对一个区块进行重采样以提出新的参数值,其余区块保持不变。
- 这确保了当前与提议参数的似然估计仅在一个区块上存在差异,从而提高其相关性。
- 使用区块特定的粒子滤波器或无偏估计器计算似然估计的对数。
- 算法采用标准的梅特罗波利斯-黑斯廷斯接受率,并结合分块似然估计。
- 基于对数似然估计方差的理论推导,确定最优粒子数量。
实验结果
研究问题
- RQ1如何以一种受控且理论可处理的方式,提高当前与提议参数下对数似然估计之间的相关性?
- RQ2对随机数进行分块是否能带来更快的混合速度并降低伪边缘MCMC算法中的方差?
- RQ3对于给定模型,分块伪边缘后验马尔可夫链蒙特卡洛方法应使用多少最优粒子数?
- RQ4在计算效率方面,分块方法与现有相关随机数方法相比表现如何?
- RQ5在哪些场景下——如面板数据或子采样——分块方法能带来最大优势?
主要发现
- 与标准PMHM相比,分块方法显著提高了对数似然估计之间的相关性,从而降低了接受率的方差。
- 与Deligiannidis等(2015)的方法相比,理论分析更为简化,数学性质更清晰。
- 在面板数据模型中,该方法显著提升了计算速度,减少了达到收敛所需的粒子数量。
- 在子采样问题中,分块PMHM在显著减少似然评估次数的同时,仍能保持相近的精度。
- 推导出最优粒子数的选择指南,并在实践中被证明是有效的。
- 该方法天然适用于具有模块化或结构化潜变量的模型,如具有群体特异性效应的面板数据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。