[论文解读] Power of Graph-Based Two-Sample Tests
本文提出了一种用于多元数据中基于图的两样本检验的统一框架,将几何图与数据深度方法进行统一推广。通过莱·卡姆的局部渐近正态性理论推导渐近效率,表明诸如弗里德曼-拉夫斯基和K-最近邻的几何图检验在$O(N^{-1/2})$替代假设下具有零皮特曼效率,同时对基于深度的检验也进行了效率比较分析。
Testing equality of two multivariate distributions is a classical problem for which many non-parametric tests have been proposed over the years. Most of the popular tests are based either on geometric graphs constructed using inter-point distances between the observations (multivariate generalizations of the Wald-Wolfowitz's runs test) or on multivariate data-depth (generalizations of the Mann-Whitney rank test). These tests are known to be asymptotically normal under the null, and consistent against fixed alternatives. This paper introduces a general notion of graph-based two-sample tests, which includes the tests described above, and provides a unified framework for analyzing their asymptotic properties. The asymptotic efficiency of a general graph-based test is derived using Le Cam's theory of local asymptotic normality, which provides a theoretical basis for comparing the tests. As a consequence, it is shown that tests based on geometric graphs such as the Friedman-Rafsky test (1979), the test based on the $K$-nearest neighbor graph, and the cross-match test of Rosenbaum (2005) have zero asymptotic (Pitman) efficiency against $O(N^{-\frac{1}{2}})$ alternatives. Asymptotic efficiencies of the tests based on multivariate depth functions (the Liu-Singh rank sum statistic (1993)) are also derived from the proposed general theory. Applications of these results are illustrated both on synthetic and real datasets.
研究动机与目标
- 将基于几何图和数据深度的现有非参数两样本检验统一到一个统一的理论框架中。
- 在一般表述下分析基于图的两样本检验的渐近性质。
- 利用莱·卡姆的局部渐近正态性理论推导渐近效率,以比较检验性能。
- 评估基于几何图的检验(如弗里德曼-拉夫斯基、K-NN、交叉匹配)在局部替代假设下的效率。
- 将效率分析扩展至基于深度的检验,如刘-辛格秩和统计量。
提出的方法
- 提出一类基于图的两样本检验的一般类,利用基于点间距离或数据深度构造的任意图。
- 应用莱·卡姆的局部渐近正态性理论,推导检验的渐近相对效率。
- 推导在原假设和局部替代假设下检验统计量的渐近分布。
- 利用检验统计量的渐近正态性,在连续替代假设下计算皮特曼效率。
- 根据构造方法,基于图中的边数或基于深度的秩构建检验统计量。
- 通过在合成数据和真实世界数据集上的模拟和实际应用验证理论结果。
实验结果
研究问题
- RQ1如何将基于几何图的检验与基于深度的检验统一到一个单一的理论框架下?
- RQ2在局部替代假设下,基于图的两样本检验的渐近效率如何?
- RQ3为何基于几何图的检验(如弗里德曼-拉夫斯基和K-NN)在$O(N^{-1/2})$替代假设下具有零渐近效率?
- RQ4在相同的渐近框架下,基于深度的检验(如刘-辛格)的效率如何比较?
- RQ5所提出的框架能否预测在有限样本设置下基于图的检验之间的性能差异?
主要发现
- 基于几何图的检验,包括弗里德曼-拉夫斯基检验、K-最近邻图检验以及罗森鲍姆的交叉匹配检验,在$O(N^{-1/2})$替代假设下具有零渐近皮特曼效率。
- 所提出的通用框架使得对基于几何图和基于深度的两样本检验进行一致的渐近分析成为可能。
- 在相同理论框架下,基于多元数据深度的刘-辛格秩和统计量被证明具有非零渐近效率。
- 在原假设和局部替代假设下,检验统计量的渐近正态性得以建立,从而支持效率比较。
- 通过在合成数据和真实数据集上的应用验证了理论效率结果,结果与预测一致。
- 该框架为基于其渐近效率特性选择最优基于图的检验提供了理论依据。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。