Skip to main content
QUICK REVIEW

[Paper Review] Pitfalls and potentials in simulation studies: Questionable research practices in comparative simulation studies allow for spurious claims of superiority of any method

Samuel Pawel, Lucas Kook|arXiv (Cornell University)|Mar 24, 2022
Data Analysis with R5 citations
TL;DR

This paper exposes how questionable research practices (QRPs) in comparative simulation studies can falsely claim methodological superiority, even for a method with no real performance gain. Through a pre-registered simulation study, the authors demonstrate that selective reporting, parameter tuning, and outcome switching can make any method appear superior, highlighting the urgent need for pre-registration, code/data sharing, and methodological transparency in simulation research.

ABSTRACT

Comparative simulation studies are workhorse tools for benchmarking statistical methods. As with other empirical studies, the success of simulation studies hinges on the quality of their design, execution and reporting. If not conducted carefully and transparently, their conclusions may be misleading. In this paper we discuss various questionable research practices which may impact the validity of simulation studies, some of which cannot be detected or prevented by the current publication process in statistics journals. To illustrate our point, we invent a novel prediction method with no expected performance gain and benchmark it in a pre-registered comparative simulation study. We show how easy it is to make the method appear superior over well-established competitor methods if questionable research practices are employed. Finally, we provide concrete suggestions for researchers, reviewers and other academic stakeholders for improving the methodological quality of comparative simulation studies, such as pre-registering simulation protocols, incentivizing neutral simulation studies and code and data sharing.

Motivation & Objective

  • To expose how questionable research practices (QRPs) in comparative simulation studies can lead to spurious claims of methodological superiority.
  • To illustrate that even a method with no real performance gain can appear superior when QRPs like selective reporting and parameter tuning are employed.
  • To highlight the lack of methodological rigor and transparency in current simulation study reporting, which undermines reproducibility and replicability.
  • To advocate for improved standards such as pre-registration of simulation protocols, neutral benchmarking, and open sharing of code and data.
  • To raise awareness among researchers, reviewers, and publishers about the risks of flawed simulation studies in shaping scientific and clinical decision-making.

Proposed method

  • Conduct a pre-registered comparative simulation study to evaluate a novel prediction method with no expected performance gain against established methods.
  • Use a fully interacted regression model (estimand ~ 0 + method:scenario) to formally test differences in performance across methods and simulation conditions.
  • Apply a single-step multiple testing correction (Hothorn et al., 2008) to adjust p-values for pairwise comparisons within each simulation condition.
  • Determine the number of simulation replications (B = 2000) based on Monte Carlo standard error requirements, using an initial small run to estimate worst-case variance.
  • Handle convergence failures by excluding non-converging simulations and reporting failure rates, with post-hoc parameter adjustments only if failure rates exceed 10%.
  • Use performance metrics such as Brier score and AUC, with visual assessment via boxplots and violin plots, and compute summary statistics including confidence intervals.

Experimental results

Research questions

  • RQ1To what extent can questionable research practices lead to false claims of superiority for a method with no real performance advantage?
  • RQ2How do selective reporting, parameter tuning, and outcome switching distort the perceived performance of statistical methods in simulation studies?
  • RQ3What are the key methodological flaws in current simulation study design and reporting that undermine reproducibility and replicability?
  • RQ4How can pre-registration and open sharing of code and data improve the credibility and transparency of comparative simulation studies?
  • RQ5What specific practices can researchers, reviewers, and publishers adopt to reduce the risk of biased or misleading simulation study results?

Key findings

  • A novel prediction method with no expected performance gain was made to appear superior to established methods through the use of questionable research practices such as selective reporting and parameter tuning.
  • The study demonstrated that even a method with no inherent advantage can achieve statistically significant superiority claims in 12.5% of simulation conditions when QRPs are applied.
  • The number of required simulation replications was determined to be 2000 to ensure a Monte Carlo standard error below 0.0001 for the Brier score estimator.
  • Convergence failures were systematically handled by excluding non-converging simulations and reporting failure rates, with post-hoc parameter adjustments only when failure rates exceeded 10%.
  • The use of a fully interacted regression model with interaction terms between methods and simulation conditions enabled formal, statistically rigorous comparison of performance differences.
  • The study revealed that current journal policies often fail to enforce methodological rigor in simulation studies, leaving room for undetected bias due to lack of simulation protocol pre-registration and code/data sharing.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.