Skip to main content
QUICK REVIEW

[論文レビュー] Workflow Techniques for the Robust Use of Bayes Factors

Daniel J. Schad, Bruno Nicenboim|arXiv (Cornell University)|Mar 15, 2021
Explainable Artificial Intelligence (XAI)被引用数 6
ひとこと要約

本稿では、事前分布への感受性、推定の不安定性、データのばらつき、不確実性下での意思決定といった要因を踏まえ、認知科学分野におけるベイズ因子の頑健性を評価する体系的なワークフローを提案する。シミュレーションベースのキャリブレーションと感受性分析を用いて、著者らはベイズ因子が不適切な事前分布の選定、不十分な有効サンプルサイズ、モデルの誤指定によって信頼性が欠けることがあることを示し、有効性のある推論を実践的に保証するため、ユーティリティに基づく意思決定と再現可能なキャリブレーションの推奨を提唱する。

ABSTRACT

Inferences about hypotheses are ubiquitous in the cognitive sciences. Bayes factors provide one general way to compare different hypotheses by their compatibility with the observed data. Those quantifications can then also be used to choose between hypotheses. While Bayes factors provide an immediate approach to hypothesis testing, they are highly sensitive to details of the data/model assumptions. Moreover it's not clear how straightforwardly this approach can be implemented in practice, and in particular how sensitive it is to the details of the computational implementation. Here, we investigate these questions for Bayes factor analyses in the cognitive sciences. We explain the statistics underlying Bayes factors as a tool for Bayesian inferences and discuss that utility functions are needed for principled decisions on hypotheses. Next, we study how Bayes factors misbehave under different conditions. This includes a study of errors in the estimation of Bayes factors. Importantly, it is unknown whether Bayes factor estimates based on bridge sampling are unbiased for complex analyses. We are the first to use simulation-based calibration as a tool to test the accuracy of Bayes factor estimates. Moreover, we study how stable Bayes factors are against different MCMC draws. We moreover study how Bayes factors depend on variation in the data. We also look at variability of decisions based on Bayes factors and how to optimize decisions using a utility function. We outline a Bayes factor workflow that researchers can use to study whether Bayes factors are robust for their individual analysis, and we illustrate this workflow using an example from the cognitive sciences. We hope that this study will provide a workflow to test the strengths and limitations of Bayes factors as a way to quantify evidence in support of scientific hypotheses. Reproducible code is available from https://osf.io/y354c/.

研究の動機と目的

  • 異なるデータ、モデル、事前分布の仮定下での認知科学応用におけるベイズ因子の頑健性を調査すること。
  • ベイズ因子推定における不安定要因を同定すること、特に不適切な事前分布の選定と不十分なMCMC有効サンプルサイズの影響を含む。
  • シミュレーションベースのキャリブレーション(SBC)を用いて、ブリッジサンプリングおよびSavage-Dickey法による推定の正確性を評価すること。
  • 繰り返しデータ抽出および再現研究におけるベイズ因子の結果のばらつきを検討すること。
  • 研究者がベイズ因子推論および意思決定の信頼性を評価するための実用的ワークフローを構築すること、その際、ユーティリティ関数とシミュレーションベースのキャリブレーションを用いる。

提案手法

  • 既知の真のモデル下でシミュレートされたデータから事前分布を回復することで、ベイズ因子推定の正確性をテストするため、シミュレーションベースのキャリブレーション(SBC)を用いる。
  • 複雑な認知科学データのため、階層ベイズモデルをフィットし、ブリッジサンプリングを用いてベイズ因子を推定するためにRパッケージbrmsを用いる。
  • 事前分布を変化させることで感受性分析を実施し、それらがベイズ因子の安定性および結論に与える影響を評価する。
  • 事前予測および事後予測チェックを実施し、モデル仮定の妥当性とデータとの整合性を評価する。
  • 意思決定を形式化し、不確実性下での結論の頑健性を評価するために、ユーティリティ関数を適用する。
  • SBCシミュレーションを用いて意思決定をキャリブレートし、事後モデル確率が真のモデル頻度と一致するかどうかを評価する。
Figure 1: Shown are the schematic relations between the data and the model, Bayes factors, and resulting inferences and decisions. The data and the model constitute a true Bayes factor, that can be used for data informed inferences and decisions (dark red arrows). However, the true Bayes factor is u
Figure 1: Shown are the schematic relations between the data and the model, Bayes factors, and resulting inferences and decisions. The data and the model constitute a true Bayes factor, that can be used for data informed inferences and decisions (dark red arrows). However, the true Bayes factor is u

実験結果

リサーチクエスチョン

  • RQ1ブリッジサンプリングによるベイズ因子推定はどの程度正確であり、どのような条件下で真のベイズ因子を回復できないのか?
  • RQ2特に小規模〜中規模のサンプル設定において、ベイズ因子は事前分布の変化に対してどの程度感受性を示すのか?
  • RQ3被験者やアイテム効果などの要因によるデータのばらつきによって、繰り返しサンプリングにおけるベイズ因子の結果がどの程度変動するのか?
  • RQ4異なるMCMCサンプリングにおけるベイズ因子推定はどの程度安定しているのか?また、信頼できる推定に必要な有効サンプルサイズはどの程度か?
  • RQ5モデルの誤指定や不適切な事前分布の選択によって生じるベイズ因子推定のバイアスを、シミュレーションベースのキャリブレーションが検出できるか?

主な発見

  • モデルが誤指定されている場合でさえ、大きなMCMCサンプルサイズを有しても、ブリッジサンプリングによるベイズ因子推定は不正確かつ不安定であることがある。
  • SBCにより、特定のモデル構成ではSavage-Dickey法における平均事後モデル確率が不適切であることが判明し、推論が信頼できないことを示唆している。
  • ベイズ因子は事前分布の仮定に極めて感受性を示し、効果量の事前分布を弱情報的範囲内で変更しただけでも、結果が著しく変化することがわかった。
  • 再現研究では、繰り返しサンプリングにおけるベイズ因子の結果が広範にわたって変動することが示され、小〜中程度の効果量では再現性が低いことが判明した。
  • 認知科学で一般的に見られるような小効果量または低パワーの研究では、サンプルサイズが大きくても、効果量が著しく大きくならない限り、ベイズ因子は結論が出にくくなる傾向にある。
  • ユーティリティに基づく意思決定は著しく頑健性を向上させ、シミュレーションベースの方法によるキャリブレーションが行われない限り、意思決定を最適化することはできないことが示された。
Figure 2: Illustration of different types of parameters for two parameters Theta 1 and Theta 2. (a) Point hypothesis in all parameters. (b) Point hypothesis in some parameters. (c) Interval hypothesis in all parameters. (d) Interval hypothesis in some parameters. (e) Full hypothesis.
Figure 2: Illustration of different types of parameters for two parameters Theta 1 and Theta 2. (a) Point hypothesis in all parameters. (b) Point hypothesis in some parameters. (c) Interval hypothesis in all parameters. (d) Interval hypothesis in some parameters. (e) Full hypothesis.

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。