Skip to main content
QUICK REVIEW

[论文解读] Dating the Break in High-dimensional Data

Runmin Wang, Xiaofeng Shao|arXiv (Cornell University)|Feb 10, 2020
Statistical Methods and Inference参考文献 16被引用 5
一句话总结

本文提出了一种基于U-统计量的估计器,用于定位高维独立数据均值中的变化点,其效率优于最小二乘法。该方法通过渐近理论和自助抽样方法建立了渐近正态性和置信区间,在p >> n且各分量存在依赖关系时依然有效。

ABSTRACT

This paper is concerned with estimation and inference for the location of a change point in the mean of independent high-dimensional data. Our change point location estimator maximizes a new U-statistic based objective function, and its convergence rate and asymptotic distribution after suitable centering and normalization are obtained under mild assumptions. Our estimator turns out to have better efficiency as compared to the least squares based counterpart in the literature. Based on the asymptotic theory, we construct a confidence interval by plugging in consistent estimates of several quantities in the normalization. We also provide a bootstrap-based confidence interval and state its asymptotic validity under suitable conditions. Through simulation studies, we demonstrate favorable finite sample performance of the new change point location estimator as compared to its least squares based counterpart, and our bootstrap-based confidence intervals, as compared to several existing competitors. The asymptotic theory based on high-dimensional U-statistic is substantially different from those developed in the literature and is of independent interest.

研究动机与目标

  • 解决在p >> n且各分量可能存在依赖关系的高维数据中检测结构断裂的挑战。
  • 开发一种新的变化点位置估计方法,其效率优于现有基于最小二乘法的方法。
  • 在较弱的正则性条件下建立所提估计器的渐近理论,包括收敛速度和渐近分布。
  • 构建变化点位置的渐近置信区间和基于自助抽样的置信区间,确保在高维设定下的有效性。
  • 为所提方法提供理论依据,并在复杂高维数据结构中通过有限样本验证其性能。

提出的方法

  • 提出一种新的基于U-统计量的客观函数以用于最大化,从而实现变化点估计,替代传统的最小二乘法。
  • 在较弱的矩和依赖性假设下,推导U-统计量最大化器的收敛速度和渐近分布。
  • 使用归一化常数的一致估计量,构建变化点位置的渐近置信区间。
  • 提出一种适应变化幅度的自助抽样置信区间方法,并在弱依赖条件下具有理论依据。
  • 利用高维U-统计量理论推导极限分布,其与经典低维渐近框架存在显著差异。
  • 应用Hájek-Rényi型和Kolmogorov型不等式,控制高维设定下部分和与交叉项的尾部概率。

实验结果

研究问题

  • RQ1基于U-统计量的估计器在高维变化点检测中是否能实现比最小二乘法更高的效率?
  • RQ2在高维且存在依赖关系的数据假设下,变化点估计器的渐近分布是什么?
  • RQ3当p >> n且各分量存在依赖关系时,如何为变化点位置构建有效的置信区间?
  • RQ4所提出的基于自助抽样的区间在变化幅度未知的高维设定下是否保持渐近有效性?
  • RQ5与现有基于最小二乘法和CUSUM的方法相比,新方法在理论和有限样本性能方面有何优势?

主要发现

  • 所提出的基于U-统计量的估计器在某些高维模型中,其效率高于基于最小二乘法的估计器。
  • 该估计器是一致的,且在适当中心化和标准化后,收敛速度为$ O_p(n^{-1/2}) $,具有非退化的渐近分布。
  • 在较弱假设下,渐近置信区间有效,且在有限样本中表现良好,尤其在变化幅度适中时。
  • 基于自助抽样的置信区间具有渐近有效性,且在有限样本中表现优异,能自适应未知的变化大小。
  • 理论分析表明,高维U-统计量的行为与经典低维理论存在根本性差异,尤其体现在依赖结构和归一化常数方面。
  • 模拟研究证实,与最小二乘法及其他竞争方法相比,所提估计器在有限样本中表现出更优的性能,尤其在高维且存在依赖关系的设定下检测变化点时。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。