[Paper Review] Testing One Hypothesis Multiple times
This paper proposes a computationally efficient method for Testing One Hypothesis Multiple times (TOHM) under stringent significance requirements, such as p < 10⁻⁷, by combining Extreme Value Theory (EVT) with Monte Carlo simulations to approximate the global p-value for models with nuisance parameters only identifiable under the alternative. The approach enables accurate inference in bump-hunting, structural change detection, and non-nested model comparisons with as few as 100 simulations.
In applied settings, tests of hypothesis where a nuisance parameter is only identifiable under the alternative often reduces into one of Testing One Hypothesis Multiple times (TOHM). Specifically, a fine discretization of the space of the non-identifiable parameter is specified, and the null hypothesis is tested against a set of sub-alternative hypothesis, one for each point of the discretization. The resulting sub-test statistics are then combined to obtain a global p-value. In this paper, we discuss a computationally efficient inferential tool to perform TOHM under stringent significance requirements, such as those typically required in the physical sciences, (e.g., p-value $<10^{-7}$). The resulting procedure leads to a generalized approach to perform inference under non-standard conditions, including non-nested models comparisons.
Motivation & Objective
- Address the challenge of statistical inference when a nuisance parameter is only identifiable under the alternative hypothesis, a common issue in high-energy physics and signal detection.
- Overcome limitations of existing methods that require case-specific derivations, covariance estimation, or full simulation of empirical processes.
- Develop a generalizable, computationally efficient framework for global p-value approximation under stringent significance thresholds (e.g., p < 10⁻⁷).
- Enable robust inference in non-standard settings such as bump-hunting, structural change detection, and non-nested model comparisons.
- Provide practical tools for selecting thresholds and number of sub-tests to ensure inference robustness.
Proposed method
- Formulate the global test as the supremum of a stochastic process W(θ) over a parameter space Θ, with the global p-value defined as P(supθ∈Θ W(θ) > c).
- Apply Extreme Value Theory (EVT) to bound the global p-value using upcrossings of a threshold c by the process W(θ), leveraging Markov’s inequality: P(sup W(θ) > c) ≤ P(W(L) > c) + E[Nc].
- Use Monte Carlo simulations to estimate the expected number of upcrossings E[Nc] and the tail probability P(W(L) > c), enabling efficient approximation of the global p-value.
- Generalize EVT-based bounds beyond the likelihood ratio test and χ² distributions, extending applicability to non-standard test statistics.
- Integrate theoretical EVT bounds with practical simulation to balance accuracy and computational cost, achieving reliable inference with as few as 100 simulations.
- Develop graphical tools to guide threshold selection (c₀) and number of sub-tests (R) for robust inference under high significance requirements.
Experimental results
Research questions
- RQ1How can global p-values be efficiently approximated in hypothesis testing when a nuisance parameter is only identifiable under the alternative?
- RQ2What is the minimal number of Monte Carlo simulations required to achieve accurate inference under stringent significance levels (e.g., p < 10⁻⁷)?
- RQ3How does the proposed TOHM method compare to Bonferroni correction in terms of power and Type I error control in high-significance settings?
- RQ4Can EVT-based bounds be generalized beyond the likelihood ratio test and χ² distribution to accommodate complex, non-standard test statistics?
- RQ5How can practical thresholds and number of sub-tests be selected to ensure robustness and accuracy in real-world applications?
Key findings
- The proposed method achieves accurate global p-value approximation with as few as 100 Monte Carlo simulations, significantly reducing computational cost compared to full empirical process simulations.
- In Example 1 (bump-hunting), TOHM achieved a significance of 4.06σ, outperforming Bonferroni (3.69σ), demonstrating improved power under high significance thresholds.
- For the dark matter signal detection (Example 2), TOHM identified a signal at θ̃ = 27.265 GeV, close to the true value of 35 GeV, despite simplifying assumptions and lack of directional data.
- In the break-point regression model (Example 3), TOHM and MHT yielded nearly identical results (11.52σ vs. 11.43σ), confirming consistency under high significance and moderate testing load.
- The method effectively controls Type I error rates under stringent significance requirements, making it suitable for high-energy physics and other fields demanding p < 10⁻⁷.
- The graphical tools for threshold and sub-test selection enhance robustness and practical usability, especially in complex, real-world inference problems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.