[论文解读] The Likelihood Ratio Test in High-Dimensional Logistic Regression Is Asymptotically a Rescaled Chi-Square
该论文证明,在高维逻辑斯蒂回归中,当预测变量数量 $ p $ 占样本量 $ n $ 的非可忽略比例时(即 $ p/n \to \kappa < 1/2 $),似然比检验(LRT)统计量 $ 2\Lambda $ 的分布收敛于一个缩放后的 $ \chi^2_k $ 分布,其缩放因子 $ \alpha(\kappa) > 1 $,这使得威尔克斯定理失效。该缩放因子依赖于维度比 $ \kappa $,本文提出一种通过非线性方程组求解该因子的方法,从而校正了标准卡方近似,避免在高维情形下产生过于保守的 p 值。
Logistic regression is used thousands of times a day to fit data, predict future outcomes, and assess the statistical significance of explanatory variables. When used for the purpose of statistical inference, logistic models produce p-values for the regression coefficients by using an approximation to the distribution of the likelihood-ratio test. Indeed, Wilks' theorem asserts that whenever we have a fixed number $p$ of variables, twice the log-likelihood ratio (LLR) $2\\Lambda$ is distributed as a $\\chi^2_k$ variable in the limit of large sample sizes $n$; here, $k$ is the number of variables being tested. In this paper, we prove that when $p$ is not negligible compared to $n$, Wilks' theorem does not hold and that the chi-square approximation is grossly incorrect; in fact, this approximation produces p-values that are far too small (under the null hypothesis). Assume that $n$ and $p$ grow large in such a way that $p/n\ ightarrow\\kappa$ for some constant $\\kappa < 1/2$. We prove that for a class of logistic models, the LLR converges to a rescaled chi-square, namely, $2\\Lambda~\\stackrel{\\mathrm{d}}{\ ightarrow}~\\alpha(\\kappa)\\chi_k^2$, where the scaling factor $\\alpha(\\kappa)$ is greater than one as soon as the dimensionality ratio $\\kappa$ is positive. Hence, the LLR is larger than classically assumed. For instance, when $\\kappa=0.3$, $\\alpha(\\kappa)\\approx1.5$. In general, we show how to compute the scaling factor by solving a nonlinear system of two equations with two unknowns. Our mathematical arguments are involved and use techniques from approximate message passing theory, non-asymptotic random matrix theory and convex geometry. We also complement our mathematical study by showing that the new limiting distribution is accurate for finite sample sizes. Finally, all the results from this paper extend to some other regression models such as the probit regression model.
研究动机与目标
- 研究在预测变量数量 $ p $ 与样本量 $ n $ 相当的高维逻辑斯蒂回归中,威尔克斯定理的有效性。
- 确定当 $ p/n \to \kappa < 1/2 $ 时,似然比检验(LRT)统计量的渐近分布。
- 校正高维情形下用于 p 值计算的标准卡方近似,该近似会导致反保守的推断。
- 提供一种计算表征极限分布的缩放因子 $ \alpha(\kappa) $ 的方法。
提出的方法
- 使用近似消息传递(AMP)理论分析高维逻辑斯蒂模型中最大似然估计量的行为。
- 应用非渐近随机矩阵理论的工具,研究 LRT 统计量的极限分布。
- 采用凸几何技术分析参数空间和似然曲面的几何结构。
- 推导出一个包含两个未知数的非线性方程组,用于计算缩放卡方分布的缩放因子 $ \alpha(\kappa) $。
- 通过留一法分析和浓度不等式控制估计误差,并推导收敛速率。
- 通过有限样本模拟验证理论结果,表明即使在中等大小的 $ n $ 和 $ p $ 下,缩放卡方近似也具有良好的准确性。
实验结果
研究问题
- RQ1当预测变量数量 $ p $ 是样本量 $ n $ 的非可忽略比例时,威尔克斯定理在逻辑斯蒂回归中是否仍然成立?
- RQ2在高维逻辑斯蒂回归中,似然比检验统计量 $ 2\Lambda $ 的极限分布是什么?
- RQ3用于缩放卡方分布的缩放因子 $ \alpha(\kappa) $ 如何依赖于维度比 $ \kappa = p/n $?
- RQ4能否通过使用 $ \alpha(\kappa) > 1 $ 的缩放卡方分布来校正 p 值的标准卡方近似?
- RQ5所提出的极限分布对有限样本大小是否准确,而不仅限于渐近情形?
主要发现
- 似然比检验统计量 $ 2\Lambda $ 的分布收敛于 $ \alpha(\kappa)\chi^2_k $,其中当 $ \kappa > 0 $ 时恒有 $ \alpha(\kappa) > 1 $,这使得标准卡方近似失效。
- 当 $ \kappa = 0.3 $ 时,缩放因子约为 $ \alpha(0.3) \approx 1.5 $,意味着 LRT 统计量比标准卡方分布预测值大 50%。
- 缩放因子 $ \alpha(\kappa) $ 通过求解由 AMP 和随机矩阵理论导出的两个未知数的非线性方程组获得。
- 标准卡方近似导致原假设下 p 值过小,从而在高维设置中产生反保守推断。
- 有限样本模拟结果证实,缩放卡方近似即使在样本量有限时也具有良好的准确性。
- 该结果可推广至其他模型(如 probit 回归),表明该现象并非仅存在于逻辑斯蒂回归中。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。