Skip to main content
QUICK REVIEW

[论文解读] Computationally Efficient Distribution Theory for Bayesian Inference of High-Dimensional Dependent Count-Valued Data

Jonathan R. Bradley, Scott H. Holan|arXiv (Cornell University)|Dec 22, 2015
Statistical Methods and Bayesian Inference参考文献 35被引用 7
一句话总结

本文提出了一种计算高效的多元对数伽马分布,用于高维、依赖性计数数据的贝叶斯推断,通过多元时空混合效应模型(MSTM)实现可扩展的时空建模。该方法在包含数百万个观测值的数据集上实现了快速的MCMC收敛和准确的预测,其有效性通过模拟研究和对美国人口普查LEHD数据的应用得到验证。

ABSTRACT

We introduce a Bayesian approach for multivariate spatio-temporal prediction for high-dimensional count-valued data. Our primary interest is when there are possibly millions of data points referenced over different variables, geographic regions, and times. This problem requires extensive methodological advancements, as jointly modeling correlated data of this size leads to the so-called "big n problem." The computational complexity of prediction in this setting is further exacerbated by acknowledging that count-valued data are naturally non-Gaussian. Thus, we develop a new computationally efficient distribution theory for this setting. In particular, we introduce a multivariate log-gamma distribution and provide substantial theoretical development including: results regarding conditional distributions, marginal distributions, an asymptotic relationship with the multivariate normal distribution, and full-conditional distributions for a Gibbs sampler. To incorporate dependence between variables, regions, and time points, a multivariate spatio-temporal mixed effects model (MSTM) is used. The results in this manuscript are extremely general, and can be used for data that exhibit fewer sources of dependency than what we consider (e.g., multivariate, spatial-only, or spatio-temporal-only data). Hence, the implications of our modeling framework may have a large impact on the general problem of jointly modeling correlated count-valued data. We show the effectiveness of our approach through a simulation study. Additionally, we demonstrate our proposed methodology with an important application analyzing data obtained from the Longitudinal Employer-Household Dynamics (LEHD) program, which is administered by the U.S. Census Bureau.

研究动机与目标

  • 解决高维、依赖性计数数据贝叶斯推断中的“大n”问题。
  • 克服将非高斯计数数据应用于潜变量高斯过程模型时MCMC抽样计算上的不可行性。
  • 发展新的分布理论,以在大规模时空计数模型中实现吉布斯抽样所需的高效全条件分布。
  • 提供一个通用框架,适用于具有复杂依赖结构的多变量、仅空间或时空计数数据。
  • 通过美国人口普查长期雇主-住户动态(LEHD)数据的真实世界应用,展示其实际效用。

提出的方法

  • 提出多元对数伽马(MLG)分布作为非高斯、高维计数数据的共轭先验。
  • 推导出MLG框架下边际分布、条件分布和全条件分布的精确形式,以支持高效的吉布斯抽样。
  • 构建一种基于MLG分布建模潜变量过程的多元时空混合效应模型(MSTM),以捕捉变量间、空间和时间上的依赖性。
  • 通过多元对数伽马(mMLG)分布的重新参数化,确保MCMC算法中具有共轭性和计算可扩展性。
  • 对精度分量(如σK, σξ)采用离散先验,以实现MCMC中快速计算和收敛。
  • 应用sMLG(缩放多元对数伽马)规格以在模拟和真实数据设置中提升预测性能。

实验结果

研究问题

  • RQ1能否发展一种新的多元分布理论,以实现高维、依赖性计数数据的计算高效贝叶斯推断?
  • RQ2多元对数伽马分布如何支持全条件分布,从而在大规模时空模型中实现快速吉布斯抽样?
  • RQ3所提出的MSTM结合MLG先验是否在MCMC收敛性和计数数据的预测准确性方面优于标准潜变量高斯过程模型?
  • RQ4该框架能否为稀疏或缺失数据(如美国社区调查(ACS)贫困估计)生成可靠的小区估计?
  • RQ5模型设定(如sMLG与标准MLG)对预测性能和计算效率有何影响?

主要发现

  • 所提出的多元对数伽马分布能够实现全条件分布的精确计算,这对于高维设置下的高效吉布斯抽样至关重要。
  • 与标准潜变量高斯过程模型因混合不良和收敛缓慢而失败的情况相比,采用MLG先验的MSTM在MCMC抽样中实现了收敛。
  • 在模拟研究中,该方法相比MALA和LAP等其他调参策略,表现出更高的预测准确性和更快的收敛速度。
  • 在佛罗里达州ACS贫困估计的真实数据应用中,该方法实现了完整的空间覆盖,95%可信区间与官方三年期估计处于同一数量级。
  • sMLG规格显著提升了预测性能,模拟和真实数据结果均证实了这一点。
  • 每次MCMC迭代的计算成本降低至每条马尔可夫链少于51秒,使得对超过400万条观测值的数据集进行可行推断成为可能。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。