[论文解读] LOO and WAIC as Model Selection Methods for Polytomous Items
本研究评估了留一法交叉验证(LOO)和瓦塔纳贝-赤池信息准则(WAIC)作为多分类项目反应理论(IRT)模型的贝叶斯模型选择方法。通过模拟数据和真实数据,研究发现尽管所有七种方法(包括AIC、BIC、DIC等)的统计功效均很高(>0.93),但WAIC和LOO的统计功效略低于DIC,但仍表现出较强的竞争力,能够有效选择正确的多分类IRT模型。
Watanabe-Akaike information criterion (WAIC; Watanabe, 2010) and leave-one-out cross validation (LOO) are two fully Bayesian model selection methods that have been shown to perform better than other traditional information-criterion based model selection methods such as AIC, BIC, and DIC in the context of dichotomous IRT model selection. In this paper, we investigated whether such superior performances of WAIC and LOO can be generalized to scenarios of polytomous IRT model selection. Specifically, we conducted a simulation study to compare the statistical power rates of WAIC and LOO with those of AIC, BIC, AICc, SABIC, and DIC in selecting the optimal model among a group of polytomous IRT ones. We also used a real data set to demonstrate the use of LOO and WAIC for polytomous IRT model selection. The findings suggest that while all seven methods have excellent statistical power (greater than 0.93) to identify the true polytomous IRT model, WAIC and LOO seem to have slightly lower statistical power than DIC, the performance of which is marginally inferior to those of the other four frequentist methods. Keywords: polytomous IRT, Bayesian, MCMC, model comparison.
研究动机与目标
- 评估WAIC和LOO在多分类IRT模型选择中的强表现是否可推广至多分类IRT模型。
- 比较WAIC、LOO与六种经典信息准则(AIC、BIC、AICc、SABIC、DIC)在选择真实多分类IRT模型时的统计功效。
- 通过真实世界的多分类IRT数据集,展示LOO和WAIC的实际应用。
- 评估完全贝叶斯模型选择方法在复杂、多类别响应数据场景下的相对有效性。
提出的方法
- 通过蒙特卡洛模拟研究,采用多种多分类IRT模型,在受控条件下比较模型选择性能。
- 使用马尔可夫链蒙特卡洛(MCMC)方法估计所有评估模型的后验分布。
- 利用后验样本计算WAIC和LOO,以评估模型拟合度和预测准确性。
- 应用传统信息准则(AIC、BIC、AICc、SABIC、DIC)作为对比基准。
- 通过不同样本量和模型配置下的统计功效率评估模型选择性能。
- 将LOO和WAIC应用于真实数据集,以说明其在多分类IRT情境下的实际效用。
实验结果
研究问题
- RQ1WAIC和LOO在多分类IRT模型中是否能保持与在二分IRT模型中相当的强模型选择性能?
- RQ2WAIC和LOO在选择真实多分类IRT模型时的统计功效率,与AIC、BIC、AICc、SABIC和DIC相比如何?
- RQ3当应用于真实多分类IRT数据时,LOO和WAIC的实际表现如何?
- RQ4在二分类设置中表现优越的贝叶斯模型选择方法,是否也适用于更复杂的多分类响应格式?
主要发现
- 所有七种模型选择方法在识别真实多分类IRT模型时均表现出极高的统计功效,超过0.93。
- WAIC和LOO的统计功效略低于DIC,而DIC的性能则略逊于其余四种经典信息准则方法。
- 尽管统计功效略低,WAIC和LOO仍表现出高度竞争力,表明其在多分类IRT模型选择中具有强稳健性。
- 模拟结果表明,LOO和WAIC是多分类IRT框架中模型比较的可靠且有效的贝叶斯替代方法。
- 真实数据应用结果证实了LOO和WAIC在真实世界多分类IRT模型评估中的实际效用和可解释性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。