Skip to main content
QUICK REVIEW

[Paper Review] Smaller $p$-values via indirect information

Peter D. Hoff|arXiv (Cornell University)|Jul 29, 2019
Statistical Methods and Bayesian Inference22 references4 citations
TL;DR

This paper introduces FAB (Frequentist, Assisted by Bayes) p-values that leverage indirect or prior information to produce smaller, more powerful p-values than standard frequentist tests while maintaining uniformity under the null hypothesis. By using a linking model to share information across parameters—especially in multiparameter normal models—the method adaptively estimates prior distributions, yielding p-values that are uniformly distributed under the null even if the linking model is misspecified.

ABSTRACT

This article develops $p$-values for evaluating means of normal populations that make use of indirect or prior information. A $p$-value of this type is based on a biased test statistic that is optimal on average with respect to a probability distribution that encodes indirect information about the mean parameter, resulting in a smaller $p$-value if the indirect information is accurate. In a variety of multiparameter settings, we show how to adaptively estimate the indirect information for each mean parameter while still maintaining uniformity of the $p$-values under their null hypotheses. This is done using a linking model through which indirect information about the mean of one population may be obtained from the data of other populations. Importantly, the linking model does not need to be correct to maintain the uniformity of the $p$-values under their null hypotheses. This methodology is illustrated in several data analysis scenarios, including small area inference, spatially arranged populations, interactions in linear regression, and generalized linear models.

Motivation & Objective

  • To develop p-values that are uniformly distributed under the null hypothesis but more powerful than standard p-values by incorporating indirect or prior information.
  • To address the challenge of improving statistical power in multiparameter settings without violating frequentist error rate guarantees.
  • To enable adaptive use of indirect information across parameters through a linking model that does not need to be perfectly specified.
  • To provide a practical framework for constructing p-values that are smaller on average than standard p-values when indirect information is accurate.
  • To demonstrate the method's robustness and utility in real-world scenarios such as small area estimation, spatial inference, and regression interactions.

Proposed method

  • The FAB p-value is derived from a test statistic that maximizes average power with respect to a prior distribution π over the mean parameter θ.
  • For normal data Y ~ N(θ, σ²), the FAB p-value is expressed as 1 - |Φ(Y/σ + 2μσ/τ²) - Φ(-Y/σ)|, where μ and τ² are the mean and variance of π.
  • In multiparameter settings, indirect information about one parameter is estimated from data on other parameters using a linking model, such as a hierarchical or multilevel model.
  • The linking model is used to estimate the prior distribution π for each parameter, enabling adaptive FAB p-values that are smaller when the indirect information is accurate.
  • The method maintains uniform p-value distribution under the null even if the linking model is misspecified, due to the frequentist guarantee of the FAB p-value.
  • The approach is extended to cases where σ² is estimated from data, and to linear and generalized linear models via empirical Bayes estimation of the linking model.

Experimental results

Research questions

  • RQ1Can p-values be made smaller than standard frequentist p-values while preserving their uniform distribution under the null hypothesis?
  • RQ2How can indirect or prior information about multiple normal means be leveraged to improve statistical power without violating frequentist error control?
  • RQ3What is the impact of model misspecification in the linking model on the validity and performance of FAB p-values?
  • RQ4In what settings—such as small area estimation or spatially structured data—does the FAB p-value offer a meaningful advantage over standard p-values?
  • RQ5How do FAB p-values behave in dependence structures common in multiple testing scenarios, and can they be safely used in procedures requiring independence?

Key findings

  • FAB p-values are uniformly distributed on (0,1) under the null hypothesis, regardless of the prior distribution used to construct them.
  • In simulations, the fraction of FAB p-values below 0.05 was 0.05–0.06 for sample sizes n ∈ {200, 400, 800, 1600}, compared to 0.05–0.07 for Wald p-values, indicating comparable or slightly better Type I error control.
  • The FAB p-value can be as small as half the size of a standard UMPU p-value when the indirect information is accurate.
  • The probability that the FAB p-value is smaller than the UMPU p-value is at least Φ(sign(̃b_j) × θ_j), suggesting a majority of FAB p-values will be smaller when the linking model is reasonably accurate.
  • Even with a misspecified Gaussian linking model, FAB p-values still outperform UMPU p-values on average across parameters, especially when the true parameter structure is approximately normal.
  • In settings with no underlying structure (e.g., parameters centered at zero), FAB p-values converge to standard UMPU p-values, preserving validity.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.