[论文解读] What to make of non-inferiority and equivalence testing with a post-specified margin?
本文探讨了在数据收集后定义等价性边际(事后指定)时等价检验的有效性与风险,认为此类做法可能引入偏差并损害统计推断。文章主张应事先指定或采用独立于数据的边际选择,提倡将标准化效应量边际作为默认选项,并强调透明度与同行评审,以确保心理学和临床研究的可信度。
In order to determine whether or not an effect is absent based on a statistical test, the recommended frequentist tool is the equivalence test. Typically, it is expected that an appropriate equivalence margin has been specified before any data are observed. Unfortunately, this can be a difficult task. If the margin is too small, then the test's power will be substantially reduced. If the margin is too large, any claims of equivalence will be meaningless. Moreover, it remains unclear how defining the margin afterwards will bias one's results. In this short article, we consider a series of hypothetical scenarios in which the margin is defined post-hoc or is otherwise considered controversial. We also review a number of relevant, potentially problematic actual studies from clinical trials research, with the aim of motivating a critical discussion as to what is acceptable and desirable in the reporting and interpretation of equivalence tests.
研究动机与目标
- 调查在等价检验中于观察数据后定义等价性边际所涉及的统计与解释风险。
- 解决在事先指定具有临床或实际意义的等价性边际方面缺乏共识及实际困难的问题。
- 提出改进心理学和临床研究中等价检验可信度与客观性的解决方案。
- 强调边际与观察数据之间独立性的重要性,以维持第一类错误率的控制。
- 倡导由期刊审稿人或监管机构等独立第三方对等价性边际进行定义或验证,尤其是在事后指定的情况下。
提出的方法
- 分析了等价性边际事后指定的假设性及现实临床试验情景,评估其统计与解释后果。
- 回顾了监管机构(如CPMP、FDA)及统计文献中关于数据依赖性边际指定危险性的现有指南与警告。
- 提出在单位不可解释时,使用标准化效应量(如Cohen’s d = 0.2)作为默认等价性边际,以增强客观性。
- 建议在未事先指定边际时,报告置信区间作为等价检验的透明替代方案。
- 引入条件等价检验(CET)政策,即期刊编辑或审稿人在研究启动前评估边际。
- 强调需要独立、无偏的边际选择,以防止利益冲突,尤其是在研究人员有发表阳性结果激励的情况下。
实验结果
研究问题
- RQ1在观察数据后定义等价性边际会产生哪些统计与解释后果?
- RQ2事后边际指定如何影响第一类错误率及等价性声明的有效性?
- RQ3在何种情况下,使用事后指定边际的等价检验仍可被视为可解释或可辩护?
- RQ4独立监督(如期刊审稿人或监管机构)在验证等价性边际方面可发挥何种作用?
- RQ5在缺乏预先指定的临床阈值时,标准化效应量如何作为可辩护的默认边际?
主要发现
- 在数据收集后定义等价性边际可能导致结果偏差及第一类错误率膨胀,从而损害等价检验的有效性。
- 即使边际是事先指定的,若其未基于临床或实际相关性,仍可能不适当,凸显了独立论证的必要性。
- 在结果单位本身不可解释时,使用标准化效应量(如Cohen’s d = 0.2)作为默认边际可提升客观性。
- 当边际未事先指定时,排除原假设且足够狭窄的置信区间可作为等价检验的透明替代方案。
- 条件等价检验(CET)政策——即编辑或审稿人在研究启动前评估边际——为提升严谨性与可信度提供了可行路径。
- 当研究人员有发表阳性结果的激励时,事后边际指定尤其成问题,因此独立定义边际对可信度至关重要。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。