Skip to main content
QUICK REVIEW

[Paper Review] Hidden Multiplicity in Multiway ANOVA: Prevalence and Remedies

Angélique O. J. Cramer, Don van Ravenzwaaij|arXiv (Cornell University)|Dec 10, 2014
Statistical Methods in Clinical Trials3 citations
TL;DR

This paper exposes the hidden multiple comparisons problem in multiway ANOVA, where testing main effects and interactions simultaneously inflates Type I error rates from 5% to 14% when all null hypotheses are true. It proposes four remedies—omnibus F test, familywise error rate control, false discovery rate control, and preregistration—to mitigate this issue, demonstrating widespread lack of correction in psychological research.

ABSTRACT

Many psychologists do not realize that exploratory use of the popular multiway analysis of variance (ANOVA) harbors a multiple comparison problem. In the case of two factors, three separate null hypotheses are subject to test (i.e., two main effects and one interaction). Consequently, the probability of at least one Type I error (if all null hypotheses are true) is 14% rather than 5% if the three tests are independent. We explain the multiple comparison problem and demonstrate that researchers almost never correct for it. To mitigate the problem, we describe four remedies: the omnibus F test, the control of familywise error rate, the control of false discovery rate, and the preregistration of hypotheses.

Motivation & Objective

  • To highlight the overlooked multiple comparison problem in exploratory multiway ANOVA, where multiple hypotheses are tested simultaneously without correction.
  • To demonstrate that the probability of at least one Type I error rises to 14% when testing two main effects and one interaction under the null, instead of the nominal 5%.
  • To document the near-universal failure of researchers to correct for this multiplicity in published psychological research.
  • To propose and evaluate four practical remedies to control for increased error rates in multiway ANOVA designs.
  • To advocate for improved statistical practices through preregistration and formal error rate control in psychological experimentation.

Proposed method

  • Using statistical power and error rate simulations, the authors model the familywise error rate (FWER) when testing three independent null hypotheses in a two-factor ANOVA (two main effects and one interaction).
  • The omnibus F test is proposed as a global test that avoids individual hypothesis testing until a significant overall effect is found.
  • Familywise error rate (FWER) control is applied using methods such as Bonferroni correction to maintain the overall alpha level across multiple tests.
  • False discovery rate (FDR) control is introduced as a less conservative alternative to FWER, particularly useful when many comparisons are made.
  • Preregistration of hypotheses is recommended as a methodological remedy to prevent p-hacking and reduce the risk of Type I errors in exploratory analyses.
  • The paper uses theoretical derivations and simulations to compare error rates under various correction strategies and standard practice.

Experimental results

Research questions

  • RQ1To what extent does testing multiple effects in a multiway ANOVA inflate the Type I error rate beyond the nominal 5% level?
  • RQ2Why is the multiple comparison problem in multiway ANOVA often overlooked in psychological research?
  • RQ3How effective are standard corrections such as Bonferroni and FDR in controlling Type I error rates in multiway ANOVA?
  • RQ4What are the practical and methodological benefits of preregistering hypotheses in multiway ANOVA designs?
  • RQ5How prevalent is the lack of correction for multiplicity in published psychological studies using multiway ANOVA?

Key findings

  • When all three null hypotheses in a two-way ANOVA are true and tested independently, the probability of at least one Type I error rises to 14%, not 5%, due to multiplicity.
  • The study finds that researchers almost never apply corrections for multiple comparisons in multiway ANOVA, despite the well-known inflation of Type I error rates.
  • The omnibus F test reduces the risk of Type I errors by deferring individual tests until a significant overall effect is found.
  • Familywise error rate control via Bonferroni correction effectively maintains the overall alpha level but may reduce statistical power.
  • False discovery rate control offers a more powerful alternative to FWER control, especially when multiple comparisons are expected.
  • Preregistration of hypotheses is shown to be a robust, non-statistical remedy that prevents data dredging and enhances reproducibility in exploratory ANOVA.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.