[论文解读] Frequency Domain Statistical Inference for High-Dimensional Time Series
本文针对高维时间序列发展了频域统计推断方法,重点在于相干性与特别是偏相干性的稳健估计和假设检验。提出一种去偏估计量,其极限分布易于处理,从而实现对不同频率下最大相干性/偏相干性的假设检验,并在大规模多重检验中实现一致的错误发现率(FDR)控制。该方法通过模拟实验和基于EEG的脑网络连通性建模得到验证。
Analyzing time series in the frequency domain enables the development of powerful tools for investigating the second-order characteristics of multivariate processes. Parameters like the spectral density matrix and its inverse, the coherence or the partial coherence, encode comprehensively the complex linear relations between the component processes of the multivariate system. In this paper, we develop inference procedures for such parameters in a high-dimensional, time series setup. Towards this goal, we first focus on the derivation of consistent estimators of the coherence and, more importantly, of the partial coherence which possess manageable limiting distributions that are suitable for testing purposes. Statistical tests of the hypothesis that the maximum over frequencies of the coherence, respectively, of the partial coherence, do not exceed a prespecified threshold value are developed. Our approach allows for testing hypotheses for individual coherences and/or partial coherences as well as for multiple testing of large sets of such parameters. In the latter case, a consistent procedure to control the false discovery rate is developed. The finite sample performance of the inference procedures introduced is investigated by means of simulations and applications to the construction of graphical interaction models for brain connectivity based on EEG data are presented.
研究动机与目标
- 解决高维时间序列谱参数统计推断的挑战,其中标准方法因维度过高而失效。
- 在高维渐近框架下,为偏相干性构建一致估计量——这对时变过程的图模型至关重要。
- 构建关于不同频率下最大相干性或偏相干性的有效假设检验统计量,实现对线性依赖结构的全局推断。
- 为大规模多重检验中的相干性/偏相干性参数提供一致的错误发现率(FDR)控制程序。
- 通过稳健且可扩展的推断工具,实现对脑网络连通性等复杂系统(如EEG数据)的实际应用。
提出的方法
- 提出一种频域偏相干性去偏估计量,校正谱密度矩阵估计量中的偏差。
- 在高维渐近框架下推导去偏偏相干性估计量的渐近分布,确保推断的有效性。
- 基于一组频率上去偏偏相干性绝对值的最大值构造检验统计量,并采用数据驱动的阈值。
- 使用修正的卡方近似确定临界值,以考虑依赖结构和维度的影响。
- 通过估计假阳性数量的期望值,实施逐步程序以控制多重检验中的错误发现率(FDR)。
- 针对每对时间序列采用个性化核函数和预白化方法进行带宽选择,以提高高维下的估计精度。
实验结果
研究问题
- RQ1在高维时间序列设定下,能否构建偏相干性的一致且渐近正态的估计量?
- RQ2在高维系统中,如何为不同频率下的最大相干性或偏相干性构建有效的假设检验?
- RQ3所提出的推断程序在现实高维场景下的有限样本表现如何?
- RQ4能否为相干性与偏相干性参数的大规模多重检验开发一致的FDR控制程序?
- RQ5在使用EEG数据检测脑网络连通性中的真实线性依赖关系时,所提出方法的表现如何?
主要发现
- 所提出的偏相干性去偏估计量实现了稳定的极限分布,即使在高维设定下也能实现有效的渐近推断。
- 基于频率间最大偏相干性的检验在控制经验大小方面接近名义水平,且统计功效较高,尤其在样本量增大时表现更优。
- 当 δ = 0.2 时,所有模拟场景下的经验FDR均为0%,表明由于检验统计量结构的保守性,实现了稳健且有效的FDR控制。
- 所提出的检验程序的统计功效始终高于基于正则化的方法,即使在保守的FDR控制条件下亦如此。
- 在EEG数据应用中,该方法成功识别出有意义的脑网络连通性模式,展现出在神经科学中的实际应用价值。
- 该方法在各种数据生成过程中保持稳健,包括VAR(1)、VMA(5)以及高维稀疏结构,在不同样本量下均表现良好。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。