Skip to main content
QUICK REVIEW

[论文解读] Gaussian approximation for the sup-norm of high-dimensional matrix-variate U-statistics and its applications

Xiaohong Chen|arXiv (Cornell University)|Jan 31, 2016
Random Matrices and Applications参考文献 48被引用 4
一句话总结

本文在不假设数据分布结构的前提下,针对高维矩阵变量子统计量在上确界范数下提出了两步高斯近似方法。在较弱的矩条件假设下,建立了当维度 $ p $ 超过样本量 $ n $ 的高维设定中,收敛速度呈多项式衰减的结论,从而通过一种新颖的重加权自助法,实现了对协方差矩阵与等级相关系数矩阵的有效推断。

ABSTRACT

This paper studies the Gaussian approximation of high-dimensional and non-degenerate U-statistics of order two under the supremum norm. We propose a two-step Gaussian approximation procedure that does not impose structural assumptions on the data distribution. Specifically, subject to mild moment conditions on the kernel, we establish the explicit rate of convergence that decays polynomially in sample size for a high-dimensional scaling limit, where the dimension can be much larger than the sample size. We also supplement a practical Gaussian wild bootstrap method to approximate the quantiles of the maxima of centered U-statistics and prove its asymptotic validity. The wild bootstrap is demonstrated on statistical applications for high-dimensional non-Gaussian data including: (i) principled and data-dependent tuning parameter selection for regularized estimation of the covariance matrix and its related functionals; (ii) simultaneous inference for the covariance and rank correlation matrices. In particular, for the thresholded covariance matrix estimator with the bootstrap selected tuning parameter, we show that the Gaussian-like convergence rates can be achieved for heavy-tailed data, which are less conservative than those obtained by the Bonferroni technique that ignores the dependency in the underlying data distribution. In addition, we also show that even for subgaussian distributions, error bounds of the bootstrapped thresholded covariance matrix estimator can be much tighter than those of the minimax estimator with a universal threshold.

研究动机与目标

  • 研究当维度 $ p $ 随样本量 $ n $ 增长时,高维U统计量在上确界范数下的渐近行为。
  • 在不假设数据分布结构的前提下,为非退化、高维矩阵变量子统计量提供高斯近似。
  • 开发一种实用的重加权自助法,用于近似中心化U统计量最大值的分位数。
  • 在一般性分布(可能具有重尾特性)下,实现对高维协方差矩阵与等级相关系数矩阵的有效统计推断。
  • 通过考虑数据依赖性,获得比经典方法(如Bonferroni校正)更紧的误差界。

提出的方法

  • 提出两步高斯近似:首先通过Hájek投影(Hoeffding分解)将U统计量近似为其线性部分,再将该线性项近似为高斯向量。
  • 利用非退化及规范V统计量的最大矩不等式与矩界,控制近似误差。
  • 在核函数具有较弱矩条件的假设下,建立上确界范数下的多项式衰减收敛速率。
  • 引入一种重加权自助程序,用于估计中心化U统计量最大值的分位数,并证明其渐近有效性。
  • 将该自助法应用于正则化协方差估计中的数据依赖性调参选择,以及协方差与等级相关系数矩阵的联合推断。
  • 利用高维数据中的依赖结构,实现比通用阈值法或Bonferroni校正更紧的误差界。

实验结果

研究问题

  • RQ1我们能否在不假设特定数据结构的前提下,对高维矩阵变量子统计量的上确界范数实现高斯近似?
  • RQ2在 $ p \gg n $ 的情况下,且仅在最小矩条件假设下,高斯近似的收敛速率是多少?
  • RQ3重加权自助法能否在高维设定下为U统计量最大值提供渐近有效的推断?
  • RQ4与经典方法(如Bonferroni校正)相比,该方法在重尾数据下的误差界有何改进?
  • RQ5基于自助法的调参选择能否使阈值化协方差矩阵估计器在重尾分布下达到类似高斯的收敛速率?

主要发现

  • 两步高斯近似在上确界范数下实现了多项式衰减的收敛速率,即使在维度 $ p $ 快于样本量 $ n $ 的情况下,只要满足较弱的矩条件,该结论依然成立。
  • 重加权自助法在近似中心化U统计量最大值的分位数方面具有渐近有效性,从而在高维设定下实现了实际可操作的推断。
  • 对于阈值化协方差矩阵估计器,由自助法选择的调参可获得比使用通用阈值的极小化极大估计器更紧的误差界,即使在次高斯数据中亦然。
  • 该方法在重尾数据下实现了类似高斯的收敛速率,优于忽略依赖结构的Bonferroni方法所给出的保守界。
  • 该近似框架适用于有界核(如Kendall’s tau)与无界核(如样本协方差)两种情形,显著拓宽了其在协方差与等级相关系数估计中的适用范围。
  • 理论误差界通过最大矩不等式与对偶性论证推导得出,在受控矩条件下可显式给出近似误差中的常数。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。