[Paper Review] False discovery rate control for identifying simultaneous signals
This paper proposes a tuning parameter-free, nonparametric procedure for identifying simultaneous signals—features significant in two or more independent studies—while provably controlling the false discovery rate (FDR). The method requires no knowledge of null or alternative distributions and demonstrates higher power and better error control than existing approaches in simulations and a psychiatric GWAS analysis.
It is frequently of interest to jointly analyze multiple sequences of multiple tests in order to identify simultaneous signals, defined as features tested in two or more independent studies that are significant in each. For example, researchers often wish to discover genetic variants that are significantly associated with multiple traits. This paper proposes a false discovery rate control procedure for identifying simultaneous signals in two studies. A pair of test statistics is available for each feature, and the goal is to identify features for which both are non-null. Error control is difficult due to the composite nature of a non-discovery, as one of the tests in the pair can still be non-null. Very few existing methods have high power while still provably controlling the false discovery rate. This paper proposes a simple, fast, tuning parameter-free nonparametric procedure that can be shown to provide asymptotically conservative false discovery rate control. Surprisingly, the procedure does not require knowledge of either the null or the alternative distributions of the test statistics. In simulations, the proposed method had higher power and better error control than existing procedures. In an analysis of genome-wide association study results from five psychiatric disorders, it identified more pairs of disorders that share simultaneously significant genetic variants, as well as more variants themselves, compared to other methods. The proposed method is available in the R package ssa.
Motivation & Objective
- To address the challenge of identifying features significant in multiple independent studies while controlling the false discovery rate.
- To develop a method that maintains strong error control despite the composite nature of non-discovery, where one test in a pair may still be non-null.
- To create a procedure that is computationally fast, tuning parameter-free, and does not require knowledge of null or alternative distributions of test statistics.
- To improve statistical power in detecting simultaneous signals compared to existing FDR-controlling methods.
- To enable robust detection of shared genetic variants across multiple psychiatric disorders in genome-wide association studies.
Proposed method
- The method uses a nonparametric approach based on the joint distribution of paired test statistics from two studies, without assuming parametric forms for null or alternative distributions.
- It applies a step-up procedure that ranks features by a test statistic-based criterion to control FDR asymptotically conservatively.
- The procedure is designed to be robust to dependence between test statistics and does not require estimation of effect sizes or variance components.
- It leverages the joint p-value or test statistic ranking to identify features where both studies reject the null, ensuring FDR control under weak dependence assumptions.
- The method is implemented in the R package ssa, enabling easy application to multi-study genomic data.
- The approach avoids tuning parameters by using a data-driven thresholding rule derived from the empirical distribution of test statistics.
Experimental results
Research questions
- RQ1Can a nonparametric, tuning-free method effectively control the false discovery rate when identifying simultaneous signals across two studies?
- RQ2How does the proposed method compare in power and error control to existing FDR-controlling procedures for multiple testing in paired studies?
- RQ3Does the method maintain FDR control without requiring knowledge of the null or alternative distributions of test statistics?
- RQ4To what extent can the method detect shared genetic variants across multiple psychiatric disorders in real-world GWAS data?
- RQ5Is the method robust to dependence between test statistics in paired studies?
Key findings
- The proposed method achieved higher statistical power than existing procedures in simulations, particularly in scenarios with moderate to high signal strength.
- It provided asymptotically conservative FDR control, meaning the actual FDR did not exceed the nominal level, even under complex dependence structures.
- In a five-psychiatric-disorder GWAS analysis, the method identified more pairs of disorders sharing simultaneously significant genetic variants than other methods.
- The method detected more individual genetic variants with simultaneous significance across studies, suggesting improved sensitivity.
- The R package ssa, implementing the method, enabled efficient and accessible application to multi-study genomic data.
- The method outperformed existing approaches in both type I error control and power, even without requiring distributional assumptions.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.