[论文解读] Distributed Bayesian Varying Coefficient Modeling Using a Gaussian Process Prior
本文提出了一种基于高斯过程先验的分布式贝叶斯变系数模型,以实现在大规模数据集上的可扩展推断。通过将数据划分为子集,采用数据增强算法进行并行MCMC推断,并利用聚合蒙特卡洛(AMC)后验分布聚合结果,该方法在保持精确的不确定性量化和计算效率的同时,实现了极小极大最优的后验收敛速率。
Varying coefficient models (VCMs) are widely used for estimating nonlinear regression functions for functional data. Their Bayesian variants using Gaussian process priors on the functional coefficients, however, have received limited attention in massive data applications, mainly due to the prohibitively slow posterior computations using Markov chain Monte Carlo (MCMC) algorithms. We address this problem using a divide-and-conquer Bayesian approach. We first create a large number of data subsamples with much smaller sizes. Then, we formulate the VCM as a linear mixed-effects model and develop a data augmentation algorithm for obtaining MCMC draws on all the subsets in parallel. Finally, we aggregate the MCMC-based estimates of subset posteriors into a single Aggregated Monte Carlo (AMC) posterior, which is used as a computationally efficient alternative to the true posterior distribution. Theoretically, we derive minimax optimal posterior convergence rates for the AMC posteriors of both the varying coefficients and the mean regression function. We provide quantification on the orders of subset sample sizes and the number of subsets. The empirical results show that the combination schemes that satisfy our theoretical assumptions, including the AMC posterior, have better estimation performance than their main competitors across diverse simulations and in a real data analysis.
研究动机与目标
- 为解决在大规模数据集上使用高斯过程先验的变系数模型(VCMs)时,基于MCMC的贝叶斯推断计算不可行的问题。
- 开发一种可扩展的分布式贝叶斯框架,在数据划分的情况下仍能保持后验精度和不确定性量化。
- 在高斯过程先验的合理光滑性假设下,推导聚合后验分布的理论收敛速率。
- 基于底层高斯过程的光滑性,提供关于子集大小和子集数量的实用指导。
提出的方法
- 通过无放回均匀抽样将完整数据集划分为k个非重叠子集,以支持并行计算。
- 将VCM重新表述为线性混合效应模型,以促进在每个子集上使用数据增强(DA)型MCMC算法进行后验抽样。
- 采用一种DA型算法,通过修改似然函数,确保每个子集的后验是真实后验的有效近似。
- 开发一种聚合蒙特卡洛(AMC)算法,将所有子集的MCMC样本合并为一个统一且一致的后验分布。
- 理论分析推导出变系数和均值回归函数的极小极大最优后验收敛速率。
- 基于多变量高斯过程先验的光滑参数v和维度d,量化子集样本大小和子集数量的最优阶。
实验结果
研究问题
- RQ1分而治之的贝叶斯方法是否能在具有GP先验的分布式VCM中实现极小极大最优的后验收敛速率?
- RQ2如何选择子集数量和子集大小,以在计算效率和统计精度之间取得平衡?
- RQ3与竞争方法相比,AMC后验聚合方法是否能保持有效的不确定性量化并降低均方误差?
- RQ4DA型算法下子集后验近似准确性的理论依据是什么?
- RQ5AMC后验的收敛速率如何依赖于底层高斯过程的光滑性及索引空间的维度?
主要发现
- 在适当的光滑性条件下,AMC后验对变系数和均值回归函数均实现了极小极大最优的后验收敛速率。
- 理论分析表明,收敛速率阶为 $ n^{-2v/(2v+d)} $,与具有GP先验的非参数回归的极小极大速率一致。
- 实证结果表明,基于AMC的组合方案相比竞争方法,能获得更短的可信区间、更高的有效样本量和更低的均方误差。
- AMC方法在各种模拟和真实数据中均保持了更好的名义覆盖概率,证实了其鲁棒的不确定性量化能力。
- 该方法计算高效,且随数据规模增长具有良好可扩展性,可实现对大规模函数型数据的实际贝叶斯推断。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。