[论文解读] Fast and scalable non-parametric Bayesian inference for Poisson point processes
本文提出了两种快速、可扩展的非参数贝叶斯方法,用于估计在区间 [0,T] 上非齐次泊松点过程的强度函数。第一种方法对 [0,T] 上划分为 N 个区间的分段常数强度使用独立的伽马先验,从而获得后验分布的闭式解;第二种方法采用伽马马尔可夫链先验,通过吉布斯采样实现 MCMC 推断。主要贡献在于理论后验收缩率达到最优,适用于 h-霍尔德连续强度函数,且第二种方法在区间的数量选择上表现出更强的鲁棒性。
We study the problem of non-parametric Bayesian estimation of the intensity function of a Poisson point process. The observations are $n$ independent realisations of a Poisson point process on the interval $[0,T]$. We propose two related approaches. In both approaches we model the intensity function as piecewise constant on $N$ bins forming a partition of the interval $[0,T]$. In the first approach the coefficients of the intensity function are assigned independent gamma priors, leading to a closed form posterior distribution. On the theoretical side, we prove that as $n ightarrow\infty,$ the posterior asymptotically concentrates around the "true", data-generating intensity function at an optimal rate for $h$-Hölder regular intensity functions ($0 < h\leq 1$). In the second approach we employ a gamma Markov chain prior on the coefficients of the intensity function. The posterior distribution is no longer available in closed form, but inference can be performed using a straightforward version of the Gibbs sampler. Both approaches scale well with sample size, but the second is much less sensitive to the choice of $N$. Practical performance of our methods is first demonstrated via synthetic data examples. We compare our second method with other existing approaches on the UK coal mining disasters data. Furthermore, we apply it to the US mass shootings data and Donald Trump's Twitter data.
研究动机与目标
- 开发快速且可扩展的非参数贝叶斯方法,用于估计泊松点过程的强度函数。
- 确保在 h-霍尔德连续强度函数上的后验集中率具有理论最优性。
- 与现有方法相比,提高对区间数量 N 选择的鲁棒性。
- 通过高效的 MCMC 算法实现实际推断,尤其适用于大规模数据集。
提出的方法
- 将强度函数建模为在将 [0,T] 划分为 N 个区间的分段常数函数,对系数分配独立的伽马先验,以实现后验分布的闭式计算。
- 在强度系数上使用伽马马尔可夫链先验,以引入平滑性,并通过简单的吉布斯采样实现 MCMC 采样。
- 推导一种可逆跳跃 MCMC 算法,以探索不同区间数 N 的模型,使用对 N 的先验以及用于模型移动的局部提议。
- 为伽马马尔可夫链模型实现一种吉布斯采样器内嵌的梅特罗波利斯算法,对平滑参数 α 和区间系数分配先验。
- 使用边际似然近似和轨迹图评估可逆跳跃 MCMC 中的模型选择和收敛性。
- 将两种方法应用于模拟数据、英国煤矿灾难数据、美国大规模枪击事件以及唐纳德·特朗普的推文数据,以展示其实际性能。
实验结果
研究问题
- RQ1我们能否实现快速且可扩展的泊松点过程强度函数的非参数贝叶斯推断?
- RQ2在分段常数强度系数上使用独立伽马先验,是否能获得闭式后验分布并实现最优后验收缩率?
- RQ3与独立伽马先验相比,伽马马尔可夫链先验是否能提升对区间数 N 选择的鲁棒性?
- RQ4在具有复杂强度模式(如不连续性或振荡)的真实世界数据集上,所提出的方法表现如何?
- RQ5在使用不同先验对 N 建模时,可逆跳跃 MCMC 中的边际似然和后验模型索引分布行为如何?
主要发现
- 独立伽马先验方法可获得后验分布的闭式解,实现快速计算,并且在 h-霍尔德连续强度函数上实现理论最优的后验收缩率。
- 伽马马尔可夫链先验方法虽无法获得闭式后验,但可通过吉布斯采样实现高效推断,且对 N 的选择显著不敏感。
- 两种方法的后验收缩率均达到 h-霍尔德连续强度函数(0 < h ≤ 1)的最优率,证实了理论最优性。
- 在模拟示例中,伽马马尔可夫链先验生成的后验实现更平滑,且能更好地捕捉复杂强度模式,包括不连续性。
- 在英国煤矿灾难数据上,第二种方法优于现有方法,尤其在边界区域和高变异性区域表现更优。
- 数值实验表明,使用独立伽马先验与可逆跳跃 MCMC 时,边际对数似然存在多个局部最大值,表明在模型空间探索中存在挑战。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。