[Paper Review] Replicability Across Multiple Studies
This paper proposes statistical methods to assess replicability across multiple studies, focusing on identifying findings consistently supported across independent investigations rather than relying on single-study significance. It introduces p-value-based procedures for both single-feature and high-dimensional settings, enabling researchers to distinguish true, reproducible signals from those driven by a single study’s noise or bias.
Meta-analysis is routinely performed in many scientific disciplines. This analysis is attractive since discoveries are possible even when all the individual studies are underpowered. However, the meta-analytic discoveries may be entirely driven by signal in a single study, and thus non-replicable. Although the great majority of meta-analyses carried out to date do not infer on the replicability of their findings, it is possible to do so. We provide a selective overview of analyses that can be carried out towards establishing replicability of the scientific findings. We describe methods for the setting where a single outcome is examined in multiple studies (as is common in systematic reviews of medical interventions), as well as for the setting where multiple studies each examine multiple features (as in genomics applications). We also discuss some of the current shortcomings and future directions.
Motivation & Objective
- To address the growing concern over non-replicable scientific findings in meta-analyses, especially when results are driven by a single underpowered or biased study.
- To develop statistical tools that explicitly test for replicability—i.e., whether a finding is supported across multiple independent studies—rather than just overall significance.
- To extend replicability analysis to high-dimensional settings, such as genomics, where multiple features are tested across multiple studies.
- To support researchers in moving beyond traditional meta-analysis by incorporating replicability as a core criterion for scientific discovery.
- To promote methodological rigor by enabling the use of summary statistics and collaborative designs that enhance replicability assessment in large-scale studies.
Proposed method
- Proposes a p-value-based approach to test the null hypothesis of no replicability, specifically that at least r out of n studies show significant evidence for a non-null effect.
- Introduces a composite null hypothesis test using the PC p-value to evaluate whether a finding is replicable across multiple studies, controlling for false discovery rates.
- Adapts existing multiple testing procedures to allow for flexible, study-specific test statistics and weights, enabling tailored replicability analysis in collaborative settings.
- Extends the framework to high-dimensional settings by testing for features with signals in at least r/n studies, applicable to neuroimaging and GWAS data.
- Introduces a two-team design where each team designs the multiple testing procedure for the other study based on their own data, ensuring independence and flexibility.
- Utilizes evidence factors in observational studies to decompose evidence into nearly independent components, enhancing robustness to different biases.
Experimental results
Research questions
- RQ1How can we distinguish findings that are driven by a single study from those that are consistently supported across multiple independent studies?
- RQ2What statistical methods can be used to test for replicability in meta-analyses when multiple studies examine the same outcome?
- RQ3How can replicability be assessed in high-dimensional settings, such as genomics, where thousands of features are tested across multiple studies?
- RQ4Can replicability be established using only a single study, provided data are split meaningfully by covariates like age or gender?
- RQ5How can collaborative research designs enhance the flexibility and robustness of replicability analysis without compromising statistical validity?
Key findings
- The proposed p-value-based replicability tests can effectively rule out findings driven by a single study, reducing the risk of false discoveries.
- In high-dimensional settings, the method identifies features with signals in at least r/n studies, enabling consistent detection of replicable signals across subjects or studies.
- The two-team collaborative design allows for independent, flexible, and robust replicability analysis, increasing the reliability of findings in multi-study research.
- The use of evidence factors in observational studies enhances the robustness of replicability assessment by separating sources of bias and increasing independence of information.
- Replicability analysis can be applied to summary-level data, such as GWAS consortia data, making it feasible to incorporate existing studies into replicability evaluations.
- The framework supports the use of flexible test statistics and weights, allowing researchers to incorporate prior beliefs about which hypotheses are more likely to be replicable.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.