[論文レビュー] Bayesian model assessment: Use of conditional vs marginal likelihoods
この論文は、潜在変数を含むモデルにおけるベイesianモデル評価における条件付き尤度と周辺尤度の重要な違いを明確にし、条件付きDIC/WAICが1単位を除いた交差検証(LOuO)に対応し、周辺DIC/WAICが1クラスタを除いた交差検証(LOcO)に対応することを示している。主な貢献は、予測の対象が個々の単位か上位レベルのクラスタかに基づいて適切な基準を選択するための原則的枠組みを提供することにある。
Typical Bayesian methods for models with latent variables (or random effects) involve directly sampling the latent variables along with the model parameters. In high-level software code for model definitions (using, e.g., BUGS, JAGS, Stan), the likelihood is therefore specified as conditional on the latent variables. This can lead researchers to perform model comparisons via conditional likelihoods, where the latent variables are considered model parameters. In other settings, typical model comparisons involve marginal likelihoods where the latent variables are integrated out. This distinction is often overlooked despite the fact that it can have a large impact on the comparisons of interest. In this paper, we clarify and illustrate these issues, focusing on the comparison of conditional and marginal Deviance Information Criteria (DICs) and Watanabe-Akaike Information Criteria (WAICs) in psychometric modeling. The conditional/marginal distinction corresponds to whether the model should be predictive for the that are in the data or for new (where clusters typically correspond to higher-level units like people or schools). Correspondingly, we show that marginal WAIC corresponds to leave-one-cluster out (LOcO) cross-validation, whereas conditional WAIC corresponds to leave-one-unit (LOuO). These results lead to recommendations on the general application of these criteria to models with latent variables.
研究の動機と目的
- 潜在変数を含むモデルにおけるベイesianモデル比較における、条件付き尤度と周辺尤度の間の広く知られつつもしばしば無視されがちな違いを解消すること。
- DIC や WAIC などのモデル評価基準における、条件付き尤度と周辺尤度の使用がもたらす影響を明確にすること。
- 心理測定モデルや階層モデルにおける、モデル比較基準と交差検証手法(LOuO 対 LOcO)との間の原則的つながりを確立すること。
- 研究者が個々の単位かクラスタ(例:個人、学校)を予測対象としているかに応じて、適切な基準を選択するための指針を提供すること。
提案手法
- 潜在変数を含むモデルにおいて、条件付きと周辺尤度のDeviance Information Criterion (DIC) および Watanabe-Akaike Information Criterion (WAIC) を比較する。
- 潜在変数を条件付け(パラメータとして扱う)ことと、それらを統合(周辺化)することの明確な区別を形式化する。
- 周辺WAICとleave-one-cluster-out (LOcO) 交差検証との理論的同等性、および条件付きWAICとleave-one-unit-out (LOuO) 交差検証との同等性を導出する。
- 特にランダム効果を含む階層モデルを用いて、アプローチの実用的適用を示す。
- ベイズ推論の原則に基づく枠組みを構築し、モデル比較における予測分布の役割を強調する。
- シミュレーションと実データの例を用いて、条件付き基準と周辺基準の選択がもたらす実用的影響を示す。
実験結果
リサーチクエスチョン
- RQ1ベイesianモデル評価において、条件付き尤度と周辺尤度は潜在変数をどのように扱うかにどのような違いをもたらすか?
- RQ2階層モデルにおいて、条件付きWAICとLOuO交差検証の関係は何か?
- RQ3潜在変数を含むモデルにおいて、周辺WAICとLOcO交差検証の関係は何か?
- RQ4予測の対象が個々の単位かクラスタかに応じて、条件付き基準と周辺基準のどちらがより適切か?
- RQ5心理測定的応用において、条件付き尤度と周辺尤度の選択がモデル選択の結果にどのように影響するか?
主な発見
- 条件付きWAICは、クラスタを固定したまま個々の観測値の予測を行うleave-one-unit-out (LOuO) 交差検証に対応する。
- 周辺WAICは、単位のクラスタ全体の予測を行うleave-one-cluster-out (LOcO) 交差検証に対応する。
- モデル比較に条件付き尤度を使用すると、特にクラスタレベルの予測を目的としている場合、個々の単位への過剰適合(overfitting)を引き起こすおそれがある。
- 新しいクラスタや上位レベルの単位を予測する目的の場合は、周辺尤度がより適切であり、潜在変数の不確実性を適切に反映する。
- 条件付き基準と周辺基準の違いは、特にランダム効果を含む階層モデルにおいて、モデル比較の結果に顕著な影響を与える。
- 本論文は、分析の予測対象に応じて条件付き基準か周辺基準かを選択する明確な根拠を提供する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。