Skip to main content
QUICK REVIEW

[Paper Review] Change-point detection for multivariate and non-Euclidean data with local dependency

Hao Chen|arXiv (Cornell University)|Mar 5, 2019
Bayesian Methods and Mixture Models20 references4 citations
TL;DR

This paper proposes a circular block permutation with random starting point to improve change-point detection in multivariate and non-Euclidean data with local dependence. By integrating this permutation scheme into a nonparametric graph-based scan statistic framework, the method controls type I error under local dependence while maintaining high power, and provides an analytic p-value approximation for efficient computation on long sequences.

ABSTRACT

In a sequence of multivariate observations or non-Euclidean data objects, such as networks, local dependence is common and could lead to false change-point discoveries. We propose a new way of permutation -- circular block permutation with a random starting point -- to address this problem. This permutation scheme is studied on a non-parametric change-point detection framework based on a similarity graph constructed on the observations, leading to a general framework for change-point detection for data with local dependency. Simulation studies show that this new framework retains the same level of power when there is no local dependency, while it controls type I error correctly for sequences with and without local dependency. We also derive an analytic p-value approximation under this new framework. The approximation works well for sequences with length in hundreds and above, making this approach fast-applicable for long data sequences.

Motivation & Objective

  • To address the issue of inflated type I error in change-point detection when observations exhibit local dependence.
  • To develop a general, nonparametric framework for detecting change-points in high-dimensional or non-Euclidean data, such as networks or multivariate time series.
  • To maintain statistical power under independence while correcting for false discoveries due to local dependence.
  • To derive an analytic p-value approximation that enables fast computation on long data sequences.
  • To provide a practical criterion for selecting the optimal block size in the circular block permutation scheme.

Proposed method

  • Introduces circular block permutation with random starting point to preserve local dependence structure during resampling.
  • Adapts the edge-count test within a scan statistic framework using a similarity graph constructed on pooled observations.
  • Uses the maximum scan statistic across all possible change-point locations as the test statistic under the new permutation scheme.
  • Derives an analytic approximation for the p-value of the test statistic, valid for sequences of length in the hundreds or more.
  • Employs a block size selection rule based on convergence of the maximum scan statistic across increasing block sizes.
  • Extends the framework to multiple change-points via recursive application of the single change-point method using binary or circular binary segmentation.

Experimental results

Research questions

  • RQ1How does local dependence affect the performance of existing nonparametric change-point detection methods based on permutation?
  • RQ2Can a modified permutation scheme preserve local dependence structure while maintaining correct type I error control?
  • RQ3What is the impact of block size on the accuracy and stability of the test statistic under local dependence?
  • RQ4Can an analytic p-value approximation be derived that is both accurate and computationally efficient for long sequences?
  • RQ5How can the proposed framework be extended to detect multiple change-points in complex data types?

Key findings

  • The proposed circular block permutation with random starting point correctly controls type I error under both independent and locally dependent data, unlike standard permutation methods.
  • The method maintains the same level of statistical power as the original method when observations are independent.
  • The analytic p-value approximation is accurate for sequences of length 100 or more, enabling fast computation without extensive resampling.
  • The maximum scan statistic stabilizes as block size increases, and a block size threshold can be identified using a ratio criterion (e.g., ratio > 0.99) to ensure sufficient dependence structure is captured.
  • The framework successfully detects change-points in multivariate autoregressive and network data, even with strong serial correlation (e.g., ρ = 0.2).
  • The extension to multiple change-points via recursive segmentation is feasible and effective, as demonstrated in simulation studies.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.