Skip to main content
QUICK REVIEW

[论文解读] High-dimensional Change-point Detection Using Generalized Homogeneity Metrics

Shubhadeep Chakraborty, Xianyang Zhang|arXiv (Cornell University)|May 19, 2021
Statistical Methods and Inference参考文献 5被引用 8
一句话总结

本文提出了一种基于广义同质性度量的非参数高维变点检测方法,用于检测超出均值和协方差结构的突变分布变化。该方法在希尔伯特空间中构建累积和过程,在高维中小样本(HDMSS)框架下推导其极限分布,并结合野生二元分割法实现具有严格显著性检验的多变点检测。

ABSTRACT

Change-point detection has been a classical problem in statistics and econometrics. This work focuses on the problem of detecting abrupt distributional changes in the data-generating distribution of a sequence of high-dimensional observations, beyond the first two moments. This has remained a substantially less explored problem in the existing literature, especially in the high-dimensional context, compared to detecting changes in the mean or the covariance structure. We develop a nonparametric methodology to (i) detect an unknown number of change-points in an independent sequence of high-dimensional observations and (ii) test for the significance of the estimated change-point locations. Our approach essentially rests upon nonparametric tests for the homogeneity of two high-dimensional distributions. We construct a single change-point location estimator via defining a cumulative sum process in an embedded Hilbert space. As the key theoretical innovation, we rigorously derive its limiting distribution under the high dimension medium sample size (HDMSS) framework. Subsequently we combine our statistic with the idea of wild binary segmentation to recursively estimate and test for multiple change-point locations. The superior performance of our methodology compared to other existing procedures is illustrated via extensive simulation studies as well as over stock prices data observed during the period of the Great Recession in the United States.

研究动机与目标

  • 填补在高维数据中检测一般分布变化(超出均值或协方差变化)的空白。
  • 开发一种非参数方法,能够检测高维独立观测中未知数量的变点。
  • 在高维中小样本(HDMSS)框架下,为变点估计量建立理论基础的极限分布。
  • 通过非参数两样本检验框架,实现对估计变点位置的显著性检验。
  • 克服现有方法在检测高阶矩变化时的局限性,或对已知变点数量的依赖。

提出的方法

  • 通过广义同质性度量在嵌入的希尔伯特空间中构建累积和过程,以检测单个变点。
  • 以能量距离统计量为基础,但通过改进的核方法扩展其能力,以捕捉一、二阶矩以外的同质性。
  • 在HDMSS框架下推导检验统计量的极限分布,确保在高维渐近下的理论有效性。
  • 应用野生二元分割法,通过递归分割数据并检验子区间,以检测多个变点。
  • 将同质性检验与递归划分策略相结合,实现对多个变点位置的估计,并附带显著性评估。
  • 采用双中心化距离和矩条件,控制方差并在高维设定下确保一致性。

实验结果

研究问题

  • RQ1当变化发生在均值和协方差结构之外时,非参数方法是否能够检测高维数据中的分布变化?
  • RQ2在HDMSS框架下,如何构建单个变点估计量并推导其极限分布?
  • RQ3当变点数量未知时,所提出方法在检测多个变点方面的表现如何?
  • RQ4与现有的能量距离和图基检验相比,该方法在检测高阶矩变化方面的表现如何?
  • RQ5当变化位于数据序列边缘时,该方法是否仍能保持功效和准确性?

主要发现

  • 所提出方法在模拟和真实世界数据中成功检测到高维数据中超出均值和协方差偏移的分布变化。
  • 在HDMSS框架下,变点估计量的极限分布被严格推导,为统计推断提供了理论依据。
  • 在检测高阶矩变化方面,该方法优于现有的能量距离和图基检验,尤其在变化不体现在位置或尺度时表现更优。
  • 结合所提同质性度量的野生二元分割法,可实现对多个变点的准确递归检测,并支持显著性检验。
  • 对美国大萧条时期股票价格的实证研究显示,该方法识别出与经济事件一致的有意义结构性断裂。
  • 即使变化位于数据序列的起始或末端,该方法仍保持稳健性,克服了先前图基方法的局限性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。