[论文解读] Classification accuracy as a proxy for two sample testing
本文证明,在高维设置下,分类准确率可作为两样本检验的有效且强大的代理统计量。若分类器的真实准确率超过随机水平任意固定裕量($ϵ > 0$),则基于置换检验和基于高斯近似的检验均具有一致性。令人惊讶的是,Fisher线性判别分析(LDA)相较于Hotelling’s $T^2$检验的渐近相对效率为$1/\sqrt{\pi}$,表明在平衡高斯模型下性能接近最优。
When data analysts train a classifier and check if its accuracy is significantly different from chance, they are implicitly performing a two-sample test. We investigate the statistical properties of this flexible approach in the high-dimensional setting. We prove two results that hold for all classifiers in any dimensions: if its true error remains $ε$-better than chance for some $ε>0$ as $d,n o \infty$, then (a) the permutation-based test is consistent (has power approaching to one), (b) a computationally efficient test based on a Gaussian approximation of the null distribution is also consistent. To get a finer understanding of the rates of consistency, we study a specialized setting of distinguishing Gaussians with mean-difference $δ$ and common (known or unknown) covariance $Σ$, when $d/n o c \in (0,\infty)$. We study variants of Fisher's linear discriminant analysis (LDA) such as "naive Bayes" in a nontrivial regime when $ε o 0$ (the Bayes classifier has true accuracy approaching 1/2), and contrast their power with corresponding variants of Hotelling's test. Surprisingly, the expressions for their power match exactly in terms of $n,d,δ,Σ$, and the LDA approach is only worse by a constant factor, achieving an asymptotic relative efficiency (ARE) of $1/\sqrtπ$ for balanced samples. We also extend our results to high-dimensional elliptical distributions with finite kurtosis. Other results of independent interest include minimax lower bounds, and the optimality of Hotelling's test when $d=o(n)$. Simulation results validate our theory, and we present practical takeaway messages along with natural open problems.
研究动机与目标
- 研究在高维数据中使用分类准确率作为两样本假设检验代理统计量的统计有效性与统计功效。
- 确定在样本量与维度同时增长时,基于分类器的检验在保持第一类错误控制和一致性方面的条件。
- 在高维渐近框架下,比较基于分类器的检验(如LDA、朴素贝叶斯)与经典参数检验(如Hotelling’s $T^2$)的检验功效。
- 为基于置换检验和高斯近似法的推断提供理论保证,均以分类准确率为统计量。
- 在特定分布假设下,量化基于分类器的检验相对于最优参数检验的相对效率。
提出的方法
- 提出一个通用框架,将分类准确率用作两样本检验的检验统计量,通过样本分割来估计误差率。
- 提出一种基于高斯近似的检验方法,用于分类误差的零分布,证明在较弱条件下具有渐近第一类错误控制。
- 分析分类准确率的置换检验,证明在与高斯近似方法相同的条件下具有一致性。
- 推导在高维极限($d/n \to c \in (0,\infty)$)下,LDA和朴素贝叶斯分类器的分类误差的渐近分布。
- 在具有有限峰度的多元正态分布与椭圆分布下,比较LDA和朴素贝叶斯与Hotelling’s $T^2$检验的检验功效。
- 利用极小极大下界和渐近相对效率(ARE)分析,评估方法的最优性与相对性能。
实验结果
研究问题
- RQ1在何种条件下,分类准确率可作为高维两样本检验的一致检验统计量?
- RQ2在高维渐近框架下,基于分类器的检验与经典参数检验(如Hotelling’s $T^2$)相比,其检验功效如何?
- RQ3在平衡高斯模型下,LDA与朴素贝叶斯相对于Hotelling’s $T^2$的渐近相对效率(ARE)是多少?
- RQ4能否使用分类误差零分布的高斯近似来构建有效且计算高效的检验?
- RQ5样本分割比例如何影响基于分类器的检验的实证功效?
主要发现
- 若分类器的真实准确率在$n,d \to \infty$时超过随机水平任意固定裕量$\epsilon > 0$,则置换检验与基于高斯近似的检验均具有一致性(功效$\to 1$)。
- 在均值差为$\delta$、协方差矩阵为$\Sigma$的平衡高斯模型下,LDA与Hotelling’s $T^2$检验在$n,d,\delta,\Sigma$方面的检验功效完全一致。
- LDA相对于Hotelling’s $T^2$的渐近相对效率(ARE)在平衡样本下为$1/\sqrt{1.5\pi} \approx 0.43$,但本文在正确理解结果后指出应为$1/\sqrt{\pi} \approx 0.58$。
- 基于高斯近似的检验在分类器满足较弱正则性条件时,可实现渐近第一类错误控制。
- 当置换次数随$n$和$d$充分增长时,置换检验在相同条件下具有一致性。
- 模拟结果表明,实证功效在样本分割比例均衡($\kappa = 0.5$)时达到最大,尽管在有限样本中,当$\kappa$极端时,由于正态近似效果差,会出现偏差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。