Skip to main content
QUICK REVIEW

[论文解读] Central Limit Theorems for Classical Likelihood Ratio Tests for High-Dimensional Normal Distributions

Tiefeng Jiang, Fan Yang|arXiv (Cornell University)|Jun 2, 2013
Random Matrices and Applications参考文献 36被引用 5
一句话总结

本文在高维多元正态模型中建立了经典似然比检验(LRT)统计量的中心极限定理(CLT),其中样本量 $ n $ 和维度 $ p $ 同时增长,且满足 $ p/n \to y \in (0,1] $。在此渐近框架下,LRT统计量收敛于具有显式均值和方差的正态分布,相较于传统卡方近似,显著提升了高维数据分析的准确性。

ABSTRACT

For random samples of size n obtained from p-variate normal distributions, we consider the classical likelihood ratio tests (LRT) for their means and covariance matrices in the high-dimensional setting. These test statistics have been extensively studied in multivariate analysis and their limiting distributions under the null hypothesis were proved to be chi-square distributions as n goes to infinity and p remains fixed. In this paper, we consider the high-dimensional case where both p and n go to infinity with p=n/y in (0, 1]. We prove that the likelihood ratio test statistics under this assumption will converge in distribution to normal distributions with explicit means and variances. We perform the simulation study to show that the likelihood ratio tests using our central limit theorems outperform those using the traditional chi-square approximations for analyzing high-dimensional data.

研究动机与目标

  • 将经典似然比检验(LRT)扩展至高维多元正态分布,超越固定-$ p $ 渐近框架。
  • 解决当 $ p $ 和 $ n $ 较大且数量级相近时,传统卡方近似在LRT统计量上表现不佳的问题。
  • 在高维渐近框架 $ p/n \to y \in (0,1] $ 下,为各类经典LRT推导出显式的中心极限定理(CLT)。
  • 通过模拟研究验证基于CLT的新推断方法,证明其在大小和功效方面优于卡方近似。
  • 为在 $ p $ 与 $ n $ 成比例的设定下使用LRT进行现代高维统计推断提供理论基础。

提出的方法

  • 利用随机矩阵理论和多元伽马函数的性质,推导在高维渐近框架 $ p/n \to y \in (0,1] $ 下LRT统计量的极限分布。
  • 应用复分析技术,包括对伽马函数比值的界估计,以控制对数似然比统计量的矩生成函数。
  • 通过推导显式的渐近均值和方差表达式,建立标准化LRT统计量收敛于正态分布的结论。
  • 利用Wishart分布和样本协方差矩阵的分布来建模原假设下的似然比。
  • 采用多元伽马函数对数的级数展开,以近似样本协方差矩阵的对数行列式。
  • 通过大量模拟实验验证CLT近似效果,比较其与传统卡方近似的第一类错误率和统计功效。

实验结果

研究问题

  • RQ1当 $ p/n \to y \in (0,1] $ 时,高维正态数据中球面对称性检验的似然比检验统计量的极限分布是什么?
  • RQ2在高维渐近框架下,当 $ p/n \to y \in (0,1] $ 时,均值向量相等性和协方差矩阵相等性检验的LRT极限分布如何表现?
  • RQ3在高维情形下,多元正态向量分量独立性检验的古典似然比检验是否可被正态分布准确近似?
  • RQ4在高维设定下,基于CLT的LRT相较于卡方近似,在大小和功效方面表现如何?
  • RQ5在高维渐近框架下,各类经典LRT的对数似然比统计量的显式渐近均值和方差是什么?

主要发现

  • 在高维正态模型中,球面对称性、均值相等性、协方差相等性及独立性检验的似然比检验统计量,当 $ p/n \to y \in (0,1] $ 时,依分布收敛于正态分布。
  • 这些极限正态分布具有显式的均值和方差,其值依赖于渐近比 $ y = \lim p/n $,且当 $ y \to 1 $ 时方差发散。
  • 当 $ y < 1 $ 时,极限方差为 $ \sigma^2 = -2[y + \log(1 - y)] $;当 $ y = 1 $ 时,方差趋于无穷大。
  • 模拟研究显示,CLT近似在高维数据中显著优于传统卡方近似,尤其在第一类错误控制和统计功效方面表现更优。
  • 即使在临界情形 $ p/n \to 1 $ 下,CLT近似依然有效,而卡方近似因秩亏缺陷而失效。
  • 理论结果通过复分析和多元伽马函数渐近展开的严格证明予以支持,尤其针对高维情形下伽马函数比值的分析。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。