Skip to main content
QUICK REVIEW

[Paper Review] Combining parameter values or $p$-values

L. Lyons, Emilien Chapon|arXiv (Cornell University)|Apr 18, 2017
Advanced Optimization Algorithms Research2 references3 citations
TL;DR

This paper provides a comprehensive framework for combining parameter estimates and p-values from multiple experiments, emphasizing the use of weighted averages via chi-squared minimization and the BLUE (Best Linear Unbiased Estimate) method for correlated measurements. It highlights that combining results using estimated uncertainties can lead to biased outcomes if correlations or variances are inaccurately assumed, and warns against naive combination of discrepant measurements.

ABSTRACT

We review the methods to combine several measurements, in the form of parameter values or $p$-values.

Motivation & Objective

  • To provide a rigorous statistical framework for combining multiple measurements of physical parameters when full data integration is not feasible.
  • To address the risks of bias in parameter combination when estimated uncertainties are used instead of true uncertainties, especially in Poisson-distributed data.
  • To clarify the conditions under which combined estimates may lie outside the range of individual measurements due to strong correlations.
  • To evaluate the limitations and pitfalls of combining p-values, particularly selection bias and non-uniqueness in transformation methods.
  • To advocate for joint analysis over combination, while offering practical alternatives when joint analysis is infeasible.

Proposed method

  • Uses weighted averaging via chi-squared minimization to combine uncorrelated measurements, with weights proportional to the inverse of variance ($w_i = 1/\sigma_i^2$).
  • Applies the BLUE method for correlated measurements, deriving optimal weights from the inverse of the covariance matrix to minimize combined variance.
  • Introduces an iterative correction for bias in cases where uncertainty estimates depend on the measured value (e.g., Poisson processes), adjusting weights using updated parameter estimates.
  • Proposes multiple p-value combination methods: Fisher’s method (product of p-values), Stouffer’s method (z-score averaging), and chi-squared-based combination using degrees of freedom.
  • Emphasizes that p-value combination is ill-posed without knowledge of alternative hypotheses and requires careful handling to avoid selection bias.
  • Uses profile likelihood methods to assess parameter constraints and validate combined results, especially in multi-parameter fits.

Experimental results

Research questions

  • RQ1How should multiple measurements of a single physical parameter be combined when their uncertainties are uncorrelated?
  • RQ2What are the consequences of using estimated uncertainties (e.g., from Poisson counts) instead of true uncertainties in parameter combination?
  • RQ3How can correlated measurements be optimally combined to minimize variance while preserving unbiasedness?
  • RQ4In what circumstances can the combined estimate lie outside the range of individual measurements, and why is this not necessarily invalid?
  • RQ5What are the limitations and risks of combining p-values from independent tests, and which method is most appropriate under what conditions?

Key findings

  • Combining uncorrelated measurements using inverse-variance weighting yields a combined estimate with uncertainty smaller than any individual measurement, provided the true uncertainties are known.
  • Using estimated uncertainties (e.g., $\sigma_i = \sqrt{x_i}$ in Poisson processes) can lead to downward bias in the combined result, as seen in the counter-example of $100\pm10$ and $1\pm1$ yielding $2\pm1$.
  • When correlations are strong ($\rho > \sigma_1/\sigma_2$), the combined estimate may lie outside the range of individual measurements, indicating that the combination is sensitive to correlation structure.
  • The BLUE method provides optimal weights for correlated measurements and allows decomposition of total uncertainty into contributions from individual sources.
  • For p-value combination, no single method is universally optimal; Fisher’s and Stouffer’s methods are commonly used, but selection bias and lack of uniqueness remain serious concerns.
  • The paper concludes that while combination is useful when joint analysis is not possible, it is always preferable to perform a full joint analysis to avoid loss of information and bias.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.