Skip to main content
QUICK REVIEW

[论文解读] A reckless guide to P-values: local evidence, global errors

Michael J. Lew|arXiv (Cornell University)|Sep 23, 2019
Meta-analysis and systematic reviews参考文献 59被引用 7
一句话总结

本文通过区分局部证据(数据对原假设的强度)与全局错误率(统计方法的长期错误特性),挑战了关于P值的普遍误解。文章主张,当与功效分析和重复实验结合使用时,P值依然具有重要价值,反驳了P值导致可重复性危机的说法。

ABSTRACT

This chapter demystifies P-values, hypothesis tests and significance tests, and introduces the concepts of local evidence and global error rates. The local evidence is embodied in extit{this} data and concerns the hypotheses of interest for extit{this} experiment, whereas the global error rate is a property of the statistical analysis and sampling procedure. It is shown using simple examples that local evidence and global error rates can be, and should be, considered together when making inferences. Power analysis for experimental design for hypothesis testing are explained, along with the more locally focussed expected P-values. Issues relating to multiple testing, HARKing, and P-hacking are explained, and it is shown that, in many situation, their effects on local evidence and global error rates are in conflict, a conflict that can always be overcome by a fresh dataset from replication of key experiments. Statistics is complicated, and so is science. There is no singular right way to do either, and universally acceptable compromises may not exist. Statistics offers a wide array of tools for assisting with scientific inference by calibrating uncertainty, but statistical inference is not a substitute for scientific inference. P-values are useful indices of evidence and deserve their place in the statistical toolbox of basic pharmacologists.

研究动机与目标

  • 解决科学研究所中P值被广泛误用和误解的问题,特别是在实验药理学领域。
  • 澄清局部证据(数据对特定假设的说明)与全局错误率(假设检验程序的长期错误特性)之间的区别。
  • 反驳P值本身导致可重复性危机的说法,表明其问题源于误解而非固有缺陷。
  • 主张在结合适当的实验设计、功效分析和重复实验的前提下,P值作为有用工具的价值。
  • 倡导基于科学推理的审慎统计推断,而非机械套用显著性检验,强调科学思维而非统计公式。

提出的方法

  • 区分局部证据(P值作为数据与原假设不一致程度的度量)与全局错误率(重复抽样下的第一类/第二类错误率)。
  • 使用简单示例说明局部证据与全局错误率如何产生分歧,以及为何必须一并考虑。
  • 引入期望P值作为评估备择假设下证据强度的工具。
  • 解释多重检验、HARKing(假设的提出在结果之后)和P值操纵对局部证据与全局错误率的影响。
  • 提倡通过使用新数据进行重复实验来解决局部证据与被放大的全局错误率之间的冲突。
  • 将统计推断重新定义为有原则的科学过程,而非机械地应用显著性检验。

实验结果

研究问题

  • RQ1在统计推断中,如何有意义地区分并调和局部证据与全局错误率?
  • RQ2为何尽管P值在科学研究中被广泛使用,却常常无法提供充分的证据?
  • RQ3P值操纵和多重检验在多大程度上损害了局部证据与全局错误控制?
  • RQ4当被适当地使用时,P值是否仍可作为科学推断中的有用工具,而非应被贝叶斯因子或置信区间所取代?
  • RQ5如何通过使用新数据进行重复实验,解决误导性局部证据与被放大的全局错误率之间的冲突?

主要发现

  • P值本身并非具有误导性;其误用源于对其作为原假设下数据不一致性的局部证据角色的误解,而非作为信念或真理的度量。
  • 局部证据(单次实验中的证据强度)与全局错误率(长期错误特性)虽有区别,但必须共同考虑才能实现可靠的推断。
  • P值操纵和多重检验可能同时扭曲局部证据与全局错误率,造成冲突,而唯有通过使用新数据的重复实验才能解决。
  • 期望P值为评估备择假设下的证据强度提供了有用工具,有助于改进实验设计与功效分析。
  • 可重复性危机并非主要由P值本身引起,而是源于不良的统计实践以及未能区分证据与错误率。
  • 统计推断应以科学推理和对P值等工具的有原则使用为指导,而非僵化地遵循显著性检验或统计公式。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。