[Paper Review] Testing Hypotheses in Particle Physics: Plots of $p_{0}$ Versus $p_{1}$
This paper introduces and analyzes $p_0$ versus $p_1$ plots—where $p_0$ and $p_1$ are p-values for testing null ($H_0$) and alternative ($H_1$) hypotheses—as a powerful graphical tool for understanding hypothesis testing in particle physics. It demonstrates how these plots clarify key concepts such as the $CL_s$ criterion, sensitivity (Punzi definition), likelihood ratios, misleading evidence, and paradoxes like the Jeffreys-Lindley paradox, offering both frequentist and Bayesian interpretations with practical applications in multi-channel searches.
For situations where we are trying to decide which of two hypotheses $H_{0}$ and $H_{1}$ provides a better description of some data, we discuss the usefulness of plots of $p_{0}$ versus $p_{1}$, where $p_{i}$ is the $p$-value for testing $H_{i}$. They provide an interesting way of understanding the difference between the standard way of excluding $H_{1}$ and the $CL_{s}$ approach; the Punzi definition of sensitivity; the relationship between $p$-values and likelihood ratios; and the probability of observing misleading evidence. They also help illustrate the Law of the Iterated Logarithm and the Jeffreys-Lindley paradox.
Motivation & Objective
- To provide a visual and analytical framework for understanding the interplay between $p$-values and likelihood ratios in hypothesis testing for new physics searches.
- To clarify the $CL_s$ criterion and its distinction from standard frequentist testing through graphical representation.
- To illustrate the concept of sensitivity using the Punzi definition within the context of $p_0$-$p_1$ plots.
- To examine the probability of misleading evidence and the Jeffreys-Lindley paradox using these plots.
- To offer practical solutions for constructing $p_0$-$p_1$ plots in the presence of nuisance parameters using bootstrap and Bayesian methods.
Proposed method
- Construct two-dimensional plots of $p_0$ (p-value for $H_0$) versus $p_1$ (p-value for $H_1$) to visualize evidence against both hypotheses simultaneously.
- Use the likelihood ratio as a test statistic $T$ to define regions in the $(p_0, p_1)$ plane corresponding to rejection or non-rejection of $H_0$ and $H_1$.
- Derive the joint probability distribution of $p_0$ and $p_1$ under $H_0$ and $H_1$ to compute error rates and statistical power.
- Apply parametric bootstrap techniques to estimate $p_0$ and $p_1$ when nuisance parameters are present, preserving frequentist error control.
- Use Bayesian marginalization over nuisance parameters via proper priors to define prior-predictive $p$-values $p_i^*$ and Bayes factors as alternatives to frequentist $p$-values.
- Plot Gaussian likelihood-ratio contours on $p_0$-$p_1$ diagrams to summarize evidence across multiple search channels in large-sample approximations.
Experimental results
Research questions
- RQ1How do $p_0$-$p_1$ plots clarify the $CL_s$ criterion for excluding new physics hypotheses?
- RQ2In what way do these plots illustrate the difference between $p$-values and likelihood ratios in evidence assessment?
- RQ3How can $p_0$-$p_1$ plots be used to visualize the probability of observing misleading evidence in hypothesis testing?
- RQ4How do $p_0$-$p_1$ plots help resolve or visualize the Jeffreys-Lindley paradox?
- RQ5What methods can be used to construct $p_0$-$p_1$ plots when nuisance parameters are present?
Key findings
- The $p_0$-$p_1$ plot provides a comprehensive visual summary of evidence against both $H_0$ and $H_1$, clearly distinguishing discovery ($H_1$ rejected), exclusion ($H_1$ rejected), and non-decision regions.
- The $CL_s$ criterion for excluding $H_1$ corresponds to a specific region in the $p_0$-$p_1$ plane where $p_0$ is small and $p_1$ is large, reflecting the need for high significance in $H_0$ rejection.
- The Punzi definition of sensitivity is naturally visualized as the area in the $p_0$-$p_1$ plane where $H_1$ is rejected under its true distribution, corresponding to high statistical power.
- The probability of observing misleading evidence—where $p_0$ is small but $p_1$ is large—is quantified by the area in the $p_0$-$p_1$ plane under $H_0$, illustrating the risk of false discovery.
- The Jeffreys-Lindley paradox is clearly illustrated in the $p_0$-$p_1$ plot, where $p$-values can lead to contradictory conclusions compared to likelihood ratios, especially in large samples.
- For nuisance parameters, Bayesian marginalization or parametric bootstrap methods allow construction of $p_0$-$p_1$ plots, with the former being computationally simpler and the latter preserving frequentist error properties.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.