Skip to main content
QUICK REVIEW

[Paper Review] On a curious bias arising when the $\sqrt{χ^2/ν}$ scaling prescription is first applied to a sub-sample of the individual results

G. D’Agostini|arXiv (Cornell University)|Jan 17, 2020
Statistical Methods and Inference7 references4 citations
TL;DR

This paper identifies a previously unrecognized bias introduced when the √(χ²/ν) scaling prescription is applied sequentially—first to a sub-sample and then to the full sample—during the combination of independent experimental results. The bias arises because the scaling lacks statistical sufficiency, distorting the final combined result even when individual measurements are consistent and Gaussian, undermining the reliability of standard error propagation in high-energy physics and related fields.

ABSTRACT

As it is well known, the standard deviation of a weighted average depends only on the individual standard deviations, but not on the dispersion of the values around the mean. This property leads sometimes to the embarrassing situation in which the combined result 'looks' somehow at odds with the individual ones. A practical way to cure the problem is to enlarge the resulting standard deviation by the $\sqrt{χ^2/ν}$ scaling, a prescription employed with arbitrary criteria on when to apply it and which individual results to use in the combination. But the `apparent' discrepancy between the combined result and the individual ones often remains. Moreover this rule does not affect the resulting `best value', even if the pattern of the individual results is highly skewed. In addition to these reasons of dissatisfaction, shared by many practitioners, the method causes another issue, recently noted on the published measurements of the charged kaon mass. It happens in fact that, if the prescription is applied twice, i.e. first to a sub-sample of the individual results and subsequently to the entire sample, then a bias on the result of the overall combination is introduced. The reason is that the prescription does not guaranty statistical sufficiency, whose importance is reminded in this script, written with a didactic spirit, with some historical notes and with a language to which most physicists are accustomed. The conclusion contains general remarks on the effective presentation of the experimental findings and a pertinent puzzle is proposed in the Appendix.

Motivation & Objective

  • To investigate the unintended bias introduced by applying the √(χ²/ν) scaling prescription twice—first to a sub-sample, then to the full sample—during result combination.
  • To highlight the failure of the √(χ²/ν) method to preserve statistical sufficiency, which compromises the validity of combined results.
  • To argue against reducing complex likelihood information to a single best value with asymmetric errors, especially when the likelihood is non-Gaussian.
  • To advocate for reporting the full likelihood function instead of summary statistics to ensure unbiased propagation and future combination of results.
  • To warn against the use of ad hoc prescriptions like √(χ²/ν) scaling, which can distort results even when individual measurements are consistent and Gaussian.

Proposed method

  • Analyzes the mathematical structure of the √(χ²/ν) scaling prescription and demonstrates its failure to preserve statistical sufficiency when applied in sequence to sub-samples and the full sample.
  • Uses a didactic, historically informed approach to illustrate how the method distorts the final combined result despite correct weighting and error estimation.
  • Applies Bayesian network modeling to represent the full probabilistic structure of experimental results, emphasizing the role of conditional probabilities and model uncertainty.
  • Derives the correct combination rule via probability theory: multiplying individual likelihood functions to form a combined likelihood, which is the theoretically sound alternative to ad hoc scaling.
  • Proposes that the likelihood function—rather than a point estimate with asymmetric errors—should be reported to preserve information for future propagation and combination.
  • Introduces a puzzle in the appendix to illustrate the counterintuitive bias that arises even in simple, well-behaved cases with independent, Gaussian-distributed measurements.

Experimental results

Research questions

  • RQ1What happens to the combined result when the √(χ²/ν) scaling is applied first to a sub-sample and then to the full sample, even if the individual results are consistent and Gaussian?
  • RQ2Why does the √(χ²/ν) scaling prescription introduce a bias in the final combined result, despite being widely used in particle physics and experimental sciences?
  • RQ3How does the lack of statistical sufficiency in the √(χ²/ν) method affect the reliability of propagated uncertainties in downstream analyses?
  • RQ4What is the theoretically correct method for combining independent experimental results when the likelihood is non-Gaussian or the χ² is not parabolic?
  • RQ5Why is reporting the full likelihood function superior to reporting a best value with asymmetric errors, especially in the context of error propagation?

Key findings

  • Applying the √(χ²/ν) scaling prescription twice—first to a sub-sample and then to the full sample—introduces a systematic bias in the final combined result, even when all individual measurements are consistent and Gaussian.
  • The √(χ²/ν) method fails to preserve statistical sufficiency, meaning it does not retain all information from the original data, leading to distorted inference.
  • The bias persists even when the pattern of individual results is highly skewed, and the method does not affect the best estimate, making the bias undetectable by standard χ² tests.
  • Reporting the full likelihood function instead of a point estimate with asymmetric errors ensures unbiased propagation and future combination, especially when the likelihood is non-Gaussian.
  • The use of likelihood ratios or negative log-likelihoods is recommended as a principled alternative to ad hoc error scaling, particularly in negative searches or non-parabolic χ² profiles.
  • The paper’s proposed puzzle in the appendix demonstrates that even in simple cases with independent, Gaussian measurements, sequential √(χ²/ν) scaling can produce misleading results, highlighting the need for caution in standard practices.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.