Skip to main content
QUICK REVIEW

[论文解读] High-Dimensional Conditionally Gaussian State Space Models with Missing Data

Joshua C. C. Chan, Aubrey Poon|arXiv (Cornell University)|Feb 7, 2023
Bayesian Modeling and Causal Inference被引用 5
一句话总结

本文提出一种基于精度矩阵的MCMC抽样方法,用于高维条件高斯状态空间模型,处理复杂缺失数据模式,利用精度矩阵的稀疏性,一步高效抽取所有缺失观测值。该方法显著加速了包含混合频率或不平衡面板数据的大规模贝叶斯VAR模型与动态因子模型的计算效率。

ABSTRACT

We develop an efficient sampling approach for handling complex missing data patterns and a large number of missing observations in conditionally Gaussian state space models. Two important examples are dynamic factor models with unbalanced datasets and large Bayesian VARs with variables in multiple frequencies. A key insight underlying the proposed approach is that the joint distribution of the missing data conditional on the observed data is Gaussian. Moreover, the inverse covariance or precision matrix of this conditional distribution is sparse, and this special structure can be exploited to substantially speed up computations. We illustrate the methodology using two empirical applications. The first application combines quarterly, monthly and weekly data using a large Bayesian VAR to produce weekly GDP estimates. In the second application, we extract latent factors from unbalanced datasets involving over a hundred monthly variables via a dynamic factor model with stochastic volatility.

研究动机与目标

  • 解决在复杂缺失数据模式(如混合频率数据与不平衡面板)下高效估计大规模时间序列模型的挑战。
  • 克服现有基于卡尔曼滤波的方法在处理高维状态空间模型中大量缺失观测值时的计算瓶颈。
  • 开发一种模块化、通用的抽样框架,适用于条件高斯状态空间模型,包括动态因子模型与具备灵活特性的大规模贝叶斯VAR模型。
  • 通过整合实时、多频率数据源与稳健的统计建模,实现及时且全面的宏观经济分析。
  • 将完整数据模型的快速估计技术扩展至缺失数据场景,在保持模型灵活性的同时维持计算效率。

提出的方法

  • 提出一种基于精度的MCMC抽样器,一步抽取所有缺失数据,条件于观测数据,充分利用条件分布的高斯性质。
  • 利用缺失数据分布精度矩阵(即逆协方差矩阵)的稀疏性与带状结构,实现快速矩阵运算。
  • 将Chan和Jeliazkov(2009)提出的基于精度的抽样算法应用于具有缺失数据的高维状态空间模型,替代计算量大的卡尔曼滤波。
  • 使用Durbin-Koopman模拟平滑器,高效抽样完整状态路径(包括缺失观测值),适用于具有时间关联约束的模型。
  • 将状态向量设计为包含滞后状态,并采用分块结构的转移矩阵与测量矩阵,以嵌入混合频率与缺失数据约束。
  • 通过允许与任意高效抽样器集成于完整数据条件高斯状态空间模型中,确保方法的模块化,提升广泛应用潜力。
Figure 1: Computation time of obtaining 10 draws against $n^{o}$ and $n^{m}$ , the numbers of observed and partially unobserved variables, respectively, with $T=300$ and $p=5$ . The four methods are: precision-based sampler with hard inter-temporal constraints (P-hard), precision-based sampler with
Figure 1: Computation time of obtaining 10 draws against $n^{o}$ and $n^{m}$ , the numbers of observed and partially unobserved variables, respectively, with $T=300$ and $p=5$ . The four methods are: precision-based sampler with hard inter-temporal constraints (P-hard), precision-based sampler with

实验结果

研究问题

  • RQ1如何在不依赖计算密集型卡尔曼滤波的前提下,高效估计具有复杂缺失数据模式的高维状态空间模型?
  • RQ2在大规模贝叶斯VAR模型与动态因子模型中,利用缺失数据精度矩阵的稀疏性能在多大程度上提升计算速度?
  • RQ3在高维设置下,缺失数据的单步抽样方法是否在MCMC效率与收敛性方面优于顺序或基于滤波的方法?
  • RQ4在涉及混合频率数据与不平衡面板的真实宏观经济应用中,所提出方法的性能如何?
  • RQ5缺失数据模式(如参差不齐的边缘或不规则发布日程)对大规模时间序列模型的估计效率与预测准确性有何影响?

主要发现

  • 所提出的基于精度的抽样器相比标准卡尔曼滤波方法实现了显著的计算加速,尤其在缺失观测值数量较大时优势明显。
  • 条件于观测数据的缺失数据联合分布为高斯分布,其精度矩阵具有稀疏性与带状结构,支持高效的矩阵求逆与抽样。
  • 在周度GDP估算应用中,模型成功整合季度、月度与周度数据,利用大规模贝叶斯VAR模型生成高频GDP预测。
  • 对于包含超过100个每月变量的动态因子模型,该方法能有效从未平衡数据集与复杂缺失数据模式中提取潜在因子。
  • MCMC样本的不效率因子表明,缺失数据与模型参数的抽样器具有合理的混合效率,多数情况下中位数不效率因子低于10。
  • 即使在参差不齐的数据模式与非同步数据发布条件下,该方法仍保持高估计精度与鲁棒性,计算可扩展性优于传统方法。
Figure 2: Computation time of obtaining 10 draws against $T$ and $p$ , the numbers of time periods and lags, respectively, with $n^{m}=5$ and $n^{o}=10$ . The four methods are: precision-based sampler with hard inter-temporal constraints (P-hard), precision-based sampler with soft constraints (P-sof
Figure 2: Computation time of obtaining 10 draws against $T$ and $p$ , the numbers of time periods and lags, respectively, with $n^{m}=5$ and $n^{o}=10$ . The four methods are: precision-based sampler with hard inter-temporal constraints (P-hard), precision-based sampler with soft constraints (P-sof

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。