[论文解读] Temporally-Reweighted Chinese Restaurant Process Mixtures for Clustering, Imputing, and Forecasting Multivariate Time Series
该论文提出了一种非参数贝叶斯模型,采用时序加权中国餐馆过程(TRCRP)混合模型,联合聚类、填补和预测具有非平稳、稀疏动态的多变量时间序列。通过在簇内使用依赖历史的簇概率和层次先验,该方法在CDC流感数据上实现了更优的预测精度,在Gapminder宏观经济时间序列中发现了可解释的簇,且无需参数假设或人工调参。
This article proposes a Bayesian nonparametric method for forecasting, imputation, and clustering in sparsely observed, multivariate time series data. The method is appropriate for jointly modeling hundreds of time series with widely varying, non-stationary dynamics. Given a collection of $N$ time series, the Bayesian model first partitions them into independent clusters using a Chinese restaurant process prior. Within a cluster, all time series are modeled jointly using a novel "temporally-reweighted" extension of the Chinese restaurant process mixture. Markov chain Monte Carlo techniques are used to obtain samples from the posterior distribution, which are then used to form predictive inferences. We apply the technique to challenging forecasting and imputation tasks using seasonal flu data from the US Center for Disease Control and Prevention, demonstrating superior forecasting accuracy and competitive imputation accuracy as compared to multiple widely used baselines. We further show that the model discovers interpretable clusters in datasets with hundreds of time series, using macroeconomic data from the Gapminder Foundation.
研究动机与目标
- 解决高维、稀疏观测的多变量时间序列在非平稳动态下的预测、填补和聚类挑战。
- 开发一种非参数贝叶斯方法,避免状态空间模型和自回归模型中常见的定制模型设定与参数假设。
- 通过发现具有共享时间动态的潜在簇,实现对数百个时间序列的联合建模。
- 在历史信号充分的区域提升预测精度,同时在未见动态区域回归至广义先验。
- 提供统一框架,支持概率聚类、填补和预测,无需固定簇数或人工调参。
提出的方法
- 模型使用中国餐馆过程(CRP)先验,基于共享时间模式对时间序列进行聚类。
- 在每个簇内,时序加权CRP通过利用时间序列前 $p$ 个时间步的信息,重新加权簇概率,扩展了标准CRP。
- 层次扩展允许在多个时间序列间共享簇结构,从而发现具有一致动态的群体。
- 使用马尔可夫链蒙特卡洛(MCMC)技术从后验分布中抽样,以实现预测推断。
- 该方法联合建模当前时间序列及其在相同时间点的其他序列值,从而提升填补精度。
- 通过依赖概率聚合MCMC样本中的后验簇分配,以表达不确定性并提取稳定簇。
实验结果
研究问题
- RQ1非参数贝叶斯模型能否在具有非平稳、稀疏动态的多变量时间序列上,联合完成聚类、填补和预测?
- RQ2与标准参数和非参数基线相比,时序加权CRP在预测和填补精度方面有何改进?
- RQ3该模型能否在不预先指定簇数的情况下,发现有意义且可解释的高维时间序列簇?
- RQ4该模型在稀疏观测数据中,多大程度上利用了跨序列依赖关系以提升填补性能?
- RQ5在历史信号有限的区域,该模型表现如何?是否在未见动态中避免了过拟合?
主要发现
- TRCRP混合模型在预测美国10个区域的流感发病率方面,优于多种基线模型,包括Facebook Prophet、多输出高斯过程、季节性ARIMA和HDP-HMM。
- 在CDC流感数据上,该模型实现了具有竞争力的填补精度,且性能对窗口大小 $p$ 的敏感度低于预期,归因于强健的跨序列依赖关系。
- 在Gapminder GDP数据上,该模型发现了九个具有连贯时间模式的可解释国家簇,例如移动电话技术采用时间线的显著差异。
- 与使用动态时间规整的k-medoids相比,聚类结果在定性上更清晰、冗余更少,尤其在较高 $k$ 值时表现更优。
- TRCRP混合模型检测到比线性互相关更精细的依赖结构,如成对依赖概率热力图所示。
- 后验簇分配在MCMC样本中具有概率性和一致性,支持不确定性感知的聚类,且无需固定簇分配。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。