[论文解读] A Model-Centric Analysis of Openness, Replication, and Reproducibility
本文通过使用概率推理分析实验组件的开放性,提出了一种以模型为中心的可重复性理论。它表明,即使消除了p-hacking或发表偏倚等常见错误做法,可重复性依然无法保证,这是由于实验设计和开放性中存在根本性的结构性障碍,从而挑战了‘消除个体不当行为即可确保可重复性’的假设。
The literature on the reproducibility crisis presents several putative causes for the proliferation of irreproducible results, including HARKing, p-hacking and publication bias. Without a theory of reproducibility, however, it is difficult to determine whether these putative causes can explain most irreproducible results. Drawing from an historically informed conception of science that is open and collaborative, we identify the components of an idealized experiment and analyze these components as a precursor to develop such a theory. Openness, we suggest, has long been intuitively proposed as a solution to irreproducibility. However, this intuition has not been validated in a theoretical framework. Our concern is that the under-theorizing of these concepts can lead to flawed inferences about the (in)validity of experimental results or integrity of individual scientists. We use probabilistic arguments and examine how openness of experimental components relates to reproducibility of results. We show that there are some impediments to obtaining reproducible results that precede many of the causes often cited in literature on the reproducibility crisis. For example, even if erroneous practices such as HARKing, p-hacking, and publication bias were absent at the individual and system level, reproducibility may still not be guaranteed.
研究动机与目标
- 解决科学研究所缺乏可重复性理论框架的问题。
- 探究为何在消除HARKing和p-hacking等众所周知的错误做法后,不可重复的结果仍持续存在。
- 通过概率视角,研究实验组件的开放性如何影响可重复性。
- 识别在个体层面方法论错误出现之前即已存在的可重复性根本性障碍。
- 验证长期以来但理论化不足的直觉——开放性可提升可重复性
提出的方法
- 基于历史视角、开放且协作的科学观念,定义理想化实验的组成部分。
- 使用概率论论证,建立实验组件开放性与结果可重复性之间关系的模型。
- 分析设计、数据和分析组件的开放性如何影响可重复结果的可能性。
- 比较个体层面错误(如p-hacking)与系统性、结构性可重复性障碍的影响。
- 形式化开放性与可重复性之间的理论联系,以评估开放性本身是否足以解决可重复性问题。
实验结果
研究问题
- RQ1在p-hacking或发表偏倚等常见错误研究实践出现之前,是否存在预存的可重复性结构性障碍?
- RQ2实验组件的开放性在概率上如何影响结果的可重复性?
- RQ3在理论上,开放性可确保可重复性的直觉信念在多大程度上可被验证?
- RQ4是否存在实验设计或报告中的根本性缺陷,即使在无个体不当行为的情况下,也会破坏可重复性?
- RQ5以模型为中心的可重复性方法与现有对可重复性危机的解释相比有何不同,又如何加以改进?
主要发现
- 即使不存在HARKing、p-hacking和发表偏倚,由于实验设计中的结构性问题,可重复性仍可能无法实现。
- 理论分析表明,若未对组件层面的透明度进行适当建模,仅靠开放性不足以保证可重复性。
- 在个体层面错误出现之前,可重复性的根本性障碍已存在,表明需要系统性改革。
- 概率建模显示,实验组件的开放性可降低复制中的不确定性,但无法完全消除。
- 对开放性与可重复性的理论化不足,导致对科学诚信和结果有效性的错误推断。
- 必须建立正式的可重复性理论,才能超越对不可重复结果的临时性解释。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。