[论文解读] Workflow Techniques for the Robust Use of Bayes Factors
本文提出了一套系统性工作流程,用于评估贝叶斯因子在认知科学研究中的稳健性,涵盖对先验分布的敏感性、估计不稳定性、数据变异性以及不确定性下的决策问题。通过基于模拟的校准和敏感性分析,作者表明贝叶斯因子可能因先验设定不佳、有效样本量不足以及模型设定错误而不可靠,并主张采用基于效用的决策方法和可重现的校准程序,以确保实践中推断的有效性。
Inferences about hypotheses are ubiquitous in the cognitive sciences. Bayes factors provide one general way to compare different hypotheses by their compatibility with the observed data. Those quantifications can then also be used to choose between hypotheses. While Bayes factors provide an immediate approach to hypothesis testing, they are highly sensitive to details of the data/model assumptions. Moreover it's not clear how straightforwardly this approach can be implemented in practice, and in particular how sensitive it is to the details of the computational implementation. Here, we investigate these questions for Bayes factor analyses in the cognitive sciences. We explain the statistics underlying Bayes factors as a tool for Bayesian inferences and discuss that utility functions are needed for principled decisions on hypotheses. Next, we study how Bayes factors misbehave under different conditions. This includes a study of errors in the estimation of Bayes factors. Importantly, it is unknown whether Bayes factor estimates based on bridge sampling are unbiased for complex analyses. We are the first to use simulation-based calibration as a tool to test the accuracy of Bayes factor estimates. Moreover, we study how stable Bayes factors are against different MCMC draws. We moreover study how Bayes factors depend on variation in the data. We also look at variability of decisions based on Bayes factors and how to optimize decisions using a utility function. We outline a Bayes factor workflow that researchers can use to study whether Bayes factors are robust for their individual analysis, and we illustrate this workflow using an example from the cognitive sciences. We hope that this study will provide a workflow to test the strengths and limitations of Bayes factors as a way to quantify evidence in support of scientific hypotheses. Reproducible code is available from https://osf.io/y354c/.
研究动机与目标
- 研究在不同数据、模型和先验假设下,贝叶斯因子在认知科学应用中的稳健性。
- 识别贝叶斯因子估计中不稳定的来源,包括先验设定不佳和马尔可夫链蒙特卡洛(MCMC)有效样本量不足。
- 使用基于模拟的校准(SBC)评估桥接抽样和Savage-Dickey方法估计的准确性。
- 考察在重复抽样和复制研究中贝叶斯因子结果的变异性。
- 为研究人员开发一种实用工作流程,通过效用函数和基于模拟的校准,评估贝叶斯因子推断与决策的可靠性。
提出的方法
- 采用基于模拟的校准(SBC)通过在已知真实模型下从模拟数据中恢复先验分布,来检验贝叶斯因子估计的准确性。
- 使用R语言中的brms包拟合分层贝叶斯模型,并通过桥接抽样方法估计复杂认知科学数据的贝叶斯因子。
- 通过改变先验分布进行先验敏感性分析,评估其对贝叶斯因子稳定性和结论的影响。
- 进行先验预测检查和后验预测检查,以评估模型假设和数据兼容性。
- 应用效用函数以形式化决策过程,并评估在不确定性下结论的稳健性。
- 使用SBC模拟进行决策校准,并评估后验模型概率是否与真实模型频率一致。

实验结果
研究问题
- RQ1通过桥接抽样获得的贝叶斯因子估计在多大程度上准确?在何种条件下它们无法恢复真实的贝叶斯因子?
- RQ2贝叶斯因子对先验分布的变化有多敏感,特别是在小到中等样本量设置下?
- RQ3由于数据变异性(如被试和项目效应)导致的重复抽样中,贝叶斯因子结果的变异性有多大?
- RQ4贝叶斯因子估计在不同MCMC抽样中有多稳定?可靠估计所需的最小有效样本量是多少?
- RQ5基于模拟的校准能否检测到因模型误设或不当先验选择导致的贝叶斯因子估计偏差?
主要发现
- 即使MCMC样本量较大,通过桥接抽样获得的贝叶斯因子估计也可能不准确且不稳定,尤其是在模型设定错误时。
- SBC结果表明,在某些模型配置下,Savage-Dickey方法的平均后验模型概率无效,表明推断不可靠。
- 贝叶斯因子对先验假设高度敏感,即使在弱信息先验范围内改变效应量先验,结果也会发生显著变化。
- 复制研究显示,贝叶斯因子在重复抽样中可能差异极大,表明小到中等效应量的可重复性较低。
- 在认知科学中常见的小效应量或低统计功效的研究中,即使样本量较大,贝叶斯因子也往往得出不确定结论,除非效应量显著增大。
- 基于效用的决策方法显著提升了稳健性,表明只有通过基于模拟的方法进行校准,决策才能实现最优化。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。