[Paper Review] Which Findings from the Functional Neuromaging Literature Can We Trust?
This study evaluates the trustworthiness of functional neuroimaging findings by comparing RFT-based family-wise error (FWE) corrected p-values against false discovery rate (FDR) control using nonparametric permutation tests. It finds that for strict cluster-defining thresholds (CDT = .001), RFT-FWE p-values closely approximate FDR control at α = .05, but for lenient thresholds (CDT = .01), only very small p-values (≤ .00001) ensure FDR significance, offering quantitative guidance for interpreting past fMRI literature.
In their recent "Cluster Failure" paper, Eklund and colleagues cast doubt on the accuracy of a widely used statistical test in functional neuroimaging. Here, we leverage nonparametric methods that control the false discovery rate to offer more nuanced, quantitative guidance about which findings in the existing literature can be trusted. We show that, in the task studies examined by Eklund et al., most clusters originally reported to be significant are indeed trustworthy by the false discovery rate benchmark.
Motivation & Objective
- To provide quantitative, data-driven guidance on which fMRI findings from the literature can be trusted despite known limitations of RFT-based FWE correction.
- To address the gap in practical interpretation of RFT-FWE p-values by comparing them to more robust FDR-based inference using nonparametric permutation methods.
- To evaluate whether commonly reported RFT-FWE significant clusters remain significant under FDR control, thereby assessing their reliability in terms of false positive rate.
- To offer a benchmark for researchers to calibrate skepticism toward past fMRI studies that used RFT-based cluster-level inference.
Proposed method
- The authors use nonparametric sign-flipping permutations (5,000 realizations per contrast) to generate empirical null distributions of cluster extents.
- For each cluster, uncorrected p-values are calculated based on the proportion of permutations yielding equal or larger cluster sizes.
- Benjamini-Hochberg FDR procedure is applied to the vector of uncorrected p-values at α = .05 to control the expected proportion of false discoveries.
- The analysis is conducted separately at two cluster-defining thresholds: CDT = .001 and CDT = .01, to assess sensitivity to threshold choice.
- Results compare RFT-FWE p-values to FDR-corrected significance, identifying which clusters would remain significant under FDR control.
- The method is applied to the same open-access fMRI task data used in Eklund et al.’s cluster failure study, ensuring consistency and reproducibility.
Experimental results
Research questions
- RQ1Which fMRI findings reported with RFT-based FWE correction can be considered trustworthy when evaluated using FDR control?
- RQ2How does the reliability of RFT-FWE significant clusters vary with different cluster-defining thresholds (CDT)?
- RQ3To what extent do RFT-FWE p-values approximate effective FDR control in real fMRI data?
- RQ4Can nonparametric permutation methods provide a more accurate benchmark for assessing the trustworthiness of past fMRI findings?
- RQ5What threshold of RFT-FWE p-value ensures FDR significance at α = .05 under lenient CDT conditions?
Key findings
- For CDT = .001, only one cluster that was significant under RFT-FWE (p ≤ .05) failed to achieve significance under FDR correction at α = .05, indicating strong alignment between RFT-FWE and FDR control.
- At CDT = .01, only clusters with RFT-FWE p-values ≤ .00001 remained significant under FDR correction, indicating that most RFT-FWE significant results at this threshold are not robust to FDR control.
- The results suggest that RFT-FWE correction at CDT = .001 provides a reasonable approximation of effective FDR control, supporting trust in the majority of studies using this threshold.
- For CDT = .01, the discrepancy between RFT-FWE and FDR control is substantial, indicating that many reported findings may be false positives when assessed by FDR standards.
- The study provides a quantitative, data-driven benchmark for evaluating the reliability of past fMRI studies that used RFT-based cluster-level inference.
- The findings support cautious interpretation of fMRI results, particularly for studies using lenient CDTs, and highlight the need for methodological refinement in cluster-level inference.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.