[Paper Review] Graph-Based Two-Sample Tests for Data with Repeated Observations
This paper proposes extended graph-based two-sample tests for data with repeated observations, where traditional similarity graphs (e.g., MST, MDP) are ambiguous due to ties. It introduces new test statistics—generalized, weighted, and max-type—featuring analytic p-value approximations to ensure robustness and scalability. The key contribution is a consistent, computationally efficient framework that maintains high power even when graphs are non-unique, validated on a phone-call network dataset with significant p-value stability across repeated observations.
In the regime of two-sample comparison, tests based on a graph constructed on observations by utilizing similarity information among them is gaining attention due to their flexibility and good performances for high-dimensional/non-Euclidean data. However, when there are repeated observations, these graph-based tests could be problematic as they are versatile to the choice of the similarity graph. We propose extended graph-based test statistics to resolve this problem. The analytic p-value approximations to these extended graph-based tests are derived to facilitate the application of these tests to large datasets. The new tests are illustrated in the analysis of a phone-call network dataset. All tests are implemented in an R package gTests.
Motivation & Objective
- To address the instability of graph-based two-sample tests when similarity graphs are non-unique due to repeated observations.
- To develop extended test statistics that remain consistent and powerful despite ambiguity in graph construction.
- To derive analytic p-value approximations for large-scale data applications, avoiding computationally expensive permutations.
- To ensure robustness and reliability in real-world datasets with tied distances or repeated observations, such as network or high-dimensional data.
- To provide a unified, off-the-shelf testing framework compatible with existing R package gTests.
Proposed method
- Extends the generalized edge-count test and weighted edge-count test to handle non-unique similarity graphs by defining new statistics based on within-sample and between-sample edge counts.
- Introduces a max-type test statistic that combines the generalized and weighted edge-counts to improve power under diverse alternatives.
- Derives analytic asymptotic distributions for the extended test statistics, enabling fast p-value computation without resampling.
- Applies variance-stabilizing weights in the weighted edge-count test to minimize variance and enhance sensitivity to location shifts.
- Uses the union and averaging strategies across multiple possible graphs to aggregate results and improve robustness.
- Validates the analytic p-values against permutation-based estimates, showing close agreement in simulation and real data.
Experimental results
Research questions
- RQ1How can graph-based two-sample tests be made robust when similarity graphs are non-unique due to repeated observations?
- RQ2Can analytic p-value approximations be derived for extended edge-count test statistics to enable large-scale application?
- RQ3How do the new test statistics perform in terms of power and Type I error control when the underlying graph is ambiguous?
- RQ4What is the impact of graph non-uniqueness on p-value stability in real-world datasets with tied distances?
- RQ5How do the extended tests compare to existing methods in detecting distributional differences in high-dimensional or network data with repeated observations?
Key findings
- The generalized and weighted edge-count tests become unstable under non-unique graphs, with p-values varying drastically across different valid graphs—e.g., from 0.004 to 0.142 in the phone-call network data.
- The proposed extended test statistics, including the max-type and union/averaging strategies, yield consistent p-values across different graph realizations, resolving instability.
- Analytic p-values derived from asymptotic distributions closely match permutation-based p-values (e.g., 0.040 vs. 0.042 for S(a)), validating their accuracy for sample sizes in the hundreds.
- The extended weighted edge-count test (Rw(a)) achieved a p-value of 0.007, rejecting the null at α=0.05, while the original edge-count test failed (p=0.040) due to variance boosting.
- The max-type test M(a)(κ) with κ=1.31 had a p-value of 0.009, closely matching the weighted test, indicating strong consistency in detecting location shifts.
- The extended tests maintain high power and stability even when the similarity graph is not uniquely defined, making them suitable for real-world data with repeated observations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.