Skip to main content
QUICK REVIEW

[Paper Review] Selective Inference for Change Point Detection in Multi-dimensional Sequences

Yuta Umezu, Ichiro Takeuchi|arXiv (Cornell University)|Jun 1, 2017
Statistical Methods and Inference32 references4 citations
TL;DR

This paper proposes a selective inference framework for detecting change points (CPs) in multi-dimensional sequences that are characterized by a subset of dimensions. By treating CP detection as a selective inference problem, the authors derive exact, non-asymptotic sampling distributions that correct for selection bias in both dimension selection and time-point selection, enabling valid statistical inference with controlled Type I error rates at the nominal level.

ABSTRACT

We study the problem of detecting change points (CPs) that are characterized by a subset of dimensions in a multi-dimensional sequence. A method for detecting those CPs can be formulated as a two-stage method: one for selecting relevant dimensions, and another for selecting CPs. It has been difficult to properly control the false detection probability of these CP detection methods because selection bias in each stage must be properly corrected. Our main contribution in this paper is to formulate a CP detection problem as a selective inference problem, and show that exact (non-asymptotic) inference is possible for a class of CP detection methods. We demonstrate the performances of the proposed selective inference framework through numerical simulations and its application to our motivating medical data analysis problem.

Motivation & Objective

  • To address the challenge of detecting change points (CPs) that are relevant in only a subset of dimensions within multi-dimensional sequences.
  • To correct for selection bias arising from two-stage CP detection: first selecting relevant dimensions, then selecting time points based on aggregated scores.
  • To enable exact, non-asymptotic statistical inference on detected CPs, ensuring proper control of false detection probability at a desired significance level (e.g., α = 0.05).
  • To develop a framework applicable to real-world data, particularly in biomedical applications such as copy number variation analysis in malignant lymphoma.

Proposed method

  • Formulates the CP detection problem as a selective inference problem, where the selection of dimensions and time points is explicitly accounted for in inference.
  • Introduces a two-stage method: (1) select a subset of dimensions based on relevance, and (2) aggregate scores across selected dimensions to form a scalar test statistic.
  • Uses selective inference theory (Taylor & Tibshirani, 2015; Lee et al., 2016) to derive the exact conditional sampling distribution of the test statistic given that a CP was selected.
  • Derives the exact sampling distribution of the test statistic under selection, enabling valid p-values and confidence intervals that account for selection bias.
  • Applies the framework to both single and multiple CP detection via a local hypothesis testing approach.
  • Employs asymptotic approximations and series expansions (e.g., Taylor expansion in μ) to derive higher-order corrections to the Type I error rate, ensuring robustness.

Experimental results

Research questions

  • RQ1Can exact, non-asymptotic inference be achieved for change point detection in multi-dimensional sequences when selection occurs in both dimension and time-point stages?
  • RQ2How can selection bias from dimension selection and time-point selection be corrected simultaneously to ensure valid Type I error control?
  • RQ3What is the exact sampling distribution of the test statistic for a detected CP, conditional on the selection mechanism used in the two-stage detection process?
  • RQ4How does the proposed framework perform in terms of power and error rate control in finite-sample settings?
  • RQ5Can the framework be effectively applied to real biomedical data, such as copy number variation analysis in lymphoma patients?

Key findings

  • The proposed selective inference framework achieves exact control of the selective Type I error rate at the nominal level α, even in finite samples.
  • The framework corrects for both dimension selection and time-point selection bias, enabling valid statistical inference on detected change points.
  • The method is approximately unbiased under selective inference, with a lower bound on power derived in the selective inference sense.
  • Numerical simulations confirm that the false detection probability is properly controlled at the desired significance level.
  • Application to a lymphoma dataset demonstrates the method’s ability to detect common CPs shared across subsets of patients, with valid inference on the detected events.
  • Higher-order corrections to the Type I error rate are derived using series expansions, improving accuracy in small samples.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.