Skip to main content
QUICK REVIEW

[Paper Review] Dimension-agnostic inference using cross U-statistics

Ilmun Kim, Aaditya Ramdas|arXiv (Cornell University)|Nov 10, 2020
Statistical Methods and Inference5 citations
TL;DR

This paper introduces a dimension-agnostic inference framework using cross U-statistics that achieves Gaussian limiting distributions regardless of how dimension $d$ scales with sample size $n$. By leveraging sample splitting, self-normalization, and variational representations of test statistics, the method produces valid inference across fixed, high-, and ultra-high-dimensional regimes, with power matching that of regime-specific U-statistics up to a $\sqrt{2}$ factor.

ABSTRACT

Classical asymptotic theory for statistical inference usually involves calibrating a statistic by fixing the dimension $d$ while letting the sample size $n$ increase to infinity. Recently, much effort has been dedicated towards understanding how these methods behave in high-dimensional settings, where $d$ and $n$ both increase to infinity together. This often leads to different inference procedures, depending on the assumptions about the dimensionality, leaving the practitioner in a bind: given a dataset with 100 samples in 20 dimensions, should they calibrate by assuming $n \gg d$, or $d/n \approx 0.2$? This paper considers the goal of dimension-agnostic inference; developing methods whose validity does not depend on any assumption on $d$ versus $n$. We introduce an approach that uses variational representations of existing test statistics along with sample splitting and self-normalization to produce a refined test statistic with a Gaussian limiting distribution, regardless of how $d$ scales with $n$. The resulting statistic can be viewed as a careful modification of degenerate U-statistics, dropping diagonal blocks and retaining off-diagonal blocks. We exemplify our technique for some classical problems including one-sample mean and covariance testing, and show that our tests have minimax rate-optimal power against appropriate local alternatives. In most settings, our cross U-statistic matches the high-dimensional power of the corresponding (degenerate) U-statistic up to a $\sqrt{2}$ factor.

Motivation & Objective

  • To address the practical dilemma in statistical inference where practitioners must choose calibration methods based on assumptions about $d$ relative to $n$.
  • To develop a single inference procedure that remains valid regardless of how dimension $d$ scales with sample size $n$.
  • To achieve minimax rate-optimal power across different asymptotic regimes without prior knowledge of the scaling of $d$ and $n$.
  • To introduce a generalizable methodological framework for dimension-agnostic inference using variational representations and sample splitting.

Proposed method

  • Uses sample splitting to divide data into disjoint subsets, enabling independent estimation of conditional expectations.
  • Applies self-normalization to the U-statistic to stabilize variance and ensure convergence to a Gaussian distribution.
  • Employs variational representations of test statistics to decouple dependence on dimensionality.
  • Constructs a refined test statistic by discarding diagonal blocks and retaining only off-diagonal U-statistic components.
  • Derives asymptotic normality of the cross U-statistic under general $d_n/n$ scaling, including fixed, high-, and ultra-high-dimensional regimes.
  • Uses a combination of Markov and Chebyshev inequalities with careful concentration bounds to control type I and II errors uniformly.
Figure 1 : Pictorial illustration of the difference between the U-statistic ( 27 ) and the proposed sample-split counterpart ( 31 ) based on the same kernel $h(x,y)$ . The U-statistic is defined as the average of all pairwise distances among observations, corresponding to all elements in the $6\time
Figure 1 : Pictorial illustration of the difference between the U-statistic ( 27 ) and the proposed sample-split counterpart ( 31 ) based on the same kernel $h(x,y)$ . The U-statistic is defined as the average of all pairwise distances among observations, corresponding to all elements in the $6\time

Experimental results

Research questions

  • RQ1Can a single statistical test maintain asymptotic validity across all scaling regimes of $d_n$ relative to $n$, including fixed, high-, and ultra-high-dimensional settings?
  • RQ2Is it possible to design inference procedures whose power adapts optimally to the underlying dimensionality without prior knowledge of the scaling?
  • RQ3Can the limiting distribution of a test statistic be made approximately Gaussian regardless of the $d/n$ regime through careful construction?
  • RQ4How does the power of dimension-agnostic inference compare to that of regime-specific U-statistic methods?
  • RQ5What general methodological principles enable dimension-agnostic inference beyond specific test statistics?

Key findings

  • The proposed cross U-statistic achieves a Gaussian limiting distribution under all $d_n/n$ scaling regimes, enabling universal calibration.
  • The method ensures asymptotic validity of level $\alpha$ tests and $1-\alpha$ confidence intervals regardless of dimensionality assumptions.
  • The test maintains minimax rate-optimal power against local alternatives, matching the performance of degenerate U-statistics up to a $\sqrt{2}$ factor.
  • Type I error is uniformly controlled across all regimes, and type II error vanishes as $\min\{m_1, m_2\} \to \infty$.
  • The approach achieves uniform asymptotic validity without requiring knowledge of $d_n$ or its growth rate relative to $n$.
  • Theoretical guarantees are established via concentration inequalities and variational representations, ensuring robustness across diverse high-dimensional settings.
Figure 2 : Illustration of the data settings with a fixed conditioning set $\mathcal{X}_{2}$ (left) and an increasing conditioning set $\mathcal{X}_{2}$ (right). When $\mathcal{X}_{2}$ is fixed, the standard CLT on $\mathcal{X}_{1}$ applies conditional on $\mathcal{X}_{2}$ . However, when $\mathcal{
Figure 2 : Illustration of the data settings with a fixed conditioning set $\mathcal{X}_{2}$ (left) and an increasing conditioning set $\mathcal{X}_{2}$ (right). When $\mathcal{X}_{2}$ is fixed, the standard CLT on $\mathcal{X}_{1}$ applies conditional on $\mathcal{X}_{2}$ . However, when $\mathcal{

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.