Skip to main content
QUICK REVIEW

[论文解读] Quantifying replicability in systematic reviews: the r-value

Liat Shenhav, Ruth Heller|arXiv (Cornell University)|Jan 31, 2015
Meta-analysis and systematic reviews参考文献 7被引用 6
一句话总结

本文引入了 r-value,这是一种新颖的度量指标,用于通过衡量元分析中显著结果在多大程度上依赖于单一研究,来量化系统综述中的可重复性。r-value 是从 N 项研究中进行逐一剔除(leave-one-out)元分析所得到的最大 p 值;r-value 越小,表明效应在多項研究中越具可重复性,从而增强健康科学综述结论的可信度。

ABSTRACT

In order to assess the effect of a health care intervention, it is useful to look at an ensemble of relevant studies. The Cochrane Collaboration's admirable goal is to provide systematic reviews of all relevant clinical studies, in order to establish whether or not there is a conclusive evidence about a specific intervention. This is done mainly by conducting a meta-analysis: a statistical synthesis of results from a series of systematically collected studies. Health practitioners often interpret a significant meta-analysis summary effect as a statement that the treatment effect is consistent across a series of studies. However, the meta-analysis significance may be driven by an effect in only one of the studies. Indeed, in an analysis of two domains of Cochrane reviews we show that in a non-negligible fraction of reviews, the removal of a single study from the meta-analysis of primary endpoints makes the conclusion non-significant. Therefore, reporting the evidence towards replicability of the effect across studies in addition to the significant meta-analysis summary effect will provide credibility to the interpretation that the effect was replicated across studies. We suggest an objective, easily computed quantity, we term the r-value, that quantifies the extent of this reliance on single studies. We suggest adding the r-values to the main results and to the forest plots of systematic reviews.

研究动机与目标

  • 为应对显著元分析结果可能由单一研究驱动的风险,从而削弱可重复性和可信度。
  • 开发一种客观且可计算的度量指标,用于量化元分析中多个研究之间可重复性的证据强度。
  • 通过在标准元分析结果中补充 r-value,提升系统综述的透明度和可靠性。
  • 为在多重主要终点和多重性校正背景下解释 r-value 提供指导。
  • 倡导将 r-value 整合至系统综述的森林图和主要结果部分。

提出的方法

  • r-value 定义为从 N 项研究中进行逐一剔除元分析所获得的最大 p 值,每次分析均排除其中一项研究。
  • 对每一项研究,重新运行元分析并排除该项研究,记录所得 p 值;这些 p 值中的最大值即为 r-value。
  • r-value 衡量对无可重复性原假设的证据强度:r-value 越小,表明可重复性越强。
  • 对于多个主要终点,使用 Bonferroni 或 Benjamini-Hochberg 方法对 r-value 进行家庭错误率(FWER)或错误发现率(FDR)校正。
  • 该方法适用于固定效应和随机效应元分析模型,如附录中的模拟结果所示。
  • 建议在森林图和主要结果中包含 r-value,以表明结论在各研究间是否稳健。

实验结果

研究问题

  • RQ1在系统综述中,显著元分析结果在多大程度上由单一研究驱动?
  • RQ2如何客观量化元分析效应在多个研究之间的可重复性?
  • RQ3多重主要终点对可重复性解释有何影响,应如何调整多重性?
  • RQ4r-value 及其校正形式在模拟和真实世界系统综述中检测对单一研究过度依赖的表现如何?
  • RQ5r-value 是否能提升健康科学中系统综述结论的可信度和透明度?

主要发现

  • 在 21 份关于乳腺癌的 Cochrane 综述样本中,13 份(62%)在剔除单一研究后结果发生变化,表明其不可重复。
  • 在 6 份关于流感的更新版 Cochrane 综述中,2 份(33%)在剔除单一研究后结果发生变化,显示出较不显著但仍重要的可重复性问题。
  • r-value 在识别出结论依赖于单一研究的元分析中表现有效,尤其当 r > 0.05 时。
  • 对于多重终点,Benjamini-Hochberg 方法在 0.05 FDR 水平下对 r-value 进行校正,识别出 Review CD005211 中四个终点中的两个具有可重复性。
  • 模拟结果表明,即使使用保守的 t 检验,显著元分析结果仍可能由单一研究驱动,即使其真实效应存在,凸显了可重复性检验的必要性。
  • r-value 提供了一种稳健且客观的度量,可补充标准元分析,量化多个研究之间可重复性的证据强度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。