[Paper Review] Asymptotically Independent U-Statistics in High-Dimensional Testing
This paper proposes a family of asymptotically independent U-statistics for high-dimensional global testing, enabling adaptive combination of p-values from different orders to achieve robust power across sparse, dense, and intermediate alternatives. The key contribution is establishing asymptotic independence and normality of U-statistics of different orders under the null, allowing valid p-value combination via Fisher's method with superior power across diverse alternatives.
Many high-dimensional hypothesis tests aim to globally examine marginal or low-dimensional features of a high-dimensional joint distribution, such as testing of mean vectors, covariance matrices and regression coefficients. This paper constructs a family of U-statistics as unbiased estimators of the $\ell_p$-norms of those features. We show that under the null hypothesis, the U-statistics of different finite orders are asymptotically independent and normally distributed. Moreover, they are also asymptotically independent with the maximum-type test statistic, whose limiting distribution is an extreme value distribution. Based on the asymptotic independence property, we propose an adaptive testing procedure which combines $p$-values computed from the U-statistics of different orders. We further establish power analysis results and show that the proposed adaptive procedure maintains high power against various alternatives.
Motivation & Objective
- Address the challenge of low statistical power in high-dimensional global testing when the true alternative structure (sparse, dense, or intermediate) is unknown.
- Develop a unified framework for constructing unbiased estimators of ℓp-norms of high-dimensional parameters (e.g., mean vectors, covariance matrices, regression coefficients).
- Establish asymptotic independence and normality of U-statistics of different orders under the null hypothesis to enable valid p-value combination.
- Propose an adaptive testing procedure that combines p-values from multiple U-statistics to maintain high power across diverse alternative types.
- Provide theoretical power analysis and numerical validation showing robust performance across a wide range of alternatives.
Proposed method
- Construct U-statistics of order a as unbiased estimators of ∥E∥a_a = ∑_{l∈L} e^a_l using kernel functions of i.i.d. observations.
- Define the U-statistic of order a as U(a) = (P^n_a)^{-1} ∑_{l∈L} ∑_{1≤i1≠⋯≠ia≤n} ∏_{k=1}^a K_l(z_{i_{(k-1)γ_l+1}}, ..., z_{i_{kγ_l}}), where K_l is an unbiased kernel for parameter e_l.
- Leverage the asymptotic independence of U-statistics of different orders under the null hypothesis, proven via joint asymptotic normality.
- Combine p-values from U-statistics of multiple orders using Fisher’s method to form an adaptive test with robust power.
- Use permutation or asymptotic approximations for the limiting distribution of the maximum-type statistic (U(∞)) under the null.
- Apply the adaptive procedure to real-world problems such as two-sample mean and covariance testing, and GWAS applications.
Experimental results
Research questions
- RQ1Can U-statistics of different orders be asymptotically independent under the high-dimensional null hypothesis?
- RQ2Does combining p-values from U-statistics of multiple orders yield a testing procedure with high power across diverse alternative types (sparse, dense, intermediate)?
- RQ3How do the proposed U-statistics perform in comparison to existing sum-of-squares and maximum-type tests under faint, sparse, and moderate alternatives?
- RQ4What is the theoretical justification for the asymptotic independence of U-statistics of different orders and their limiting distributions?
- RQ5Can the adaptive procedure maintain high power when the true alternative structure is unknown or varies across settings?
Key findings
- U-statistics of different finite orders are asymptotically independent and jointly normally distributed under the null hypothesis, enabling valid p-value combination.
- The U-statistic of order a=2 is a sum-of-squares-type estimator, while higher orders (e.g., a→∞) approximate the maximum-type statistic, capturing different alternative structures.
- In the extreme faint case (Model 1), U(1) and U(2) achieve empirical power of 0.416 and 0.484 at p=300, respectively, with the adaptive procedure (adpUf2) reaching 0.781.
- In the extreme sparse case (Model 2), higher-order U-statistics (U(4), U(∞)) achieve near-1 power, while U(1) and U(2) drop to 0.049 and 0.122 at p=300.
- In the moderately sparse and faint case (Model 3), finite-order U(5) is most powerful (0.946 at p=300), and the adaptive procedure (adpUf2) maintains high power (0.940 at p=300) where other methods like Chen, Sriva, and Schott fail.
- The proposed adaptive procedure (adpUf2) achieves comparable or better power than computationally expensive methods (Pan1, Pan2) while being significantly more efficient, especially at larger p.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.