[Paper Review] Signal Significance in Particle Physics
This paper reviews the frequentist foundation of statistical significance (p-value) in particle physics, emphasizing its role in validating discoveries like the top quark and W boson. It highlights the challenges of bias in variable selection and proposes rigorous p-value estimation using multi-variate tests, concluding that a p-value of ~1.6×10⁻⁶ supports a 5σ-level discovery claim when combined across correlated observables.
The concept of the "statistical significance" of an observation, and how it is used in particle physics experiments is reviewed. More properly known as a "p-value," the statistical foundations for this concept are reviewed from a freqentist perspective. The discovery of the top quark at the Fermilab Tevatron Collider and a more recent analysis of data recorded at Fermilab are used to illustrate practical applications of these concepts.
Motivation & Objective
- To clarify the frequentist interpretation of statistical significance (p-value) in particle physics experiments.
- To analyze historical examples where significance was inconsistently quantified, such as the top quark and W boson discoveries.
- To examine the impact of a posteriori variable selection on p-value reliability in multi-variate significance testing.
- To evaluate the robustness of p-value estimation when systematic uncertainties and correlations are present.
- To advocate for independent data validation as the gold standard for confirming serendipitous discoveries.
Proposed method
- Uses frequentist hypothesis testing to define significance as the probability of observing data as extreme as, or more extreme than, the observed result under the null hypothesis.
- Applies the Kolmogorov-Smirnov (K-S) test to compare data distributions with Standard Model predictions across multiple kinematic variables.
- Combines individual p-values from 9 selected variables using a product-based formula to estimate a global p-value: $ P_{tot} = \prod_{m=1}^{9} \left[ \sum_{k=0}^{m-1} \frac{-(\ln p_{tot})^k}{k!} \right] $.
- Employs Monte Carlo simulations to estimate the sampling distribution of the K-S distance and derive p-values for goodness-of-fit tests.
- Assesses the impact of data-driven variable selection on p-value validity, acknowledging that such choices introduce bias and invalidate standard p-value interpretations.
- Recommends blind analysis and independent data sets to mitigate selection bias in discovery claims.
Experimental results
Research questions
- RQ1How is statistical significance formally defined in the frequentist framework, and what role does the p-value play in hypothesis testing?
- RQ2Why is the p-value from a multi-variate analysis unreliable when the variables are selected after observing the data?
- RQ3What is the effect of systematic uncertainties and variable correlations on the reliability of combined p-values?
- RQ4How can researchers ensure that a discovery claim is not driven by unconscious bias in data selection or analysis choices?
- RQ5What constitutes a sufficient statistical threshold for claiming a discovery in high-energy physics?
Key findings
- The p-value for the combined analysis of 13 'superjet' events in the CDF experiment was estimated at 1.6×10⁻⁶, corresponding to a 5σ significance level.
- The authors caution that this p-value is not strictly valid due to a posteriori selection of the 9 kinematic variables, which introduces selection bias.
- Despite this limitation, the p-value remains a useful indicator of significance when interpreted with caution and validated through independent data.
- The study demonstrates that multi-variate p-value combination is sensitive to correlations and systematic uncertainties, which can distort the final significance estimate.
- The paper concludes that the most reliable method for confirming a discovery remains independent replication with a new data sample of comparable statistical power.
- A p-value of approximately 10⁻⁶ is widely accepted as sufficient evidence for a discovery claim in particle physics, though this depends on rigorous methodological controls.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.