[论文解读] Assessing the Performance of Diagnostic Classification Models in Small Sample Contexts with Different Estimation Methods
本研究通过全面的模拟设计,使用最大似然(ML)、贝叶斯和非参数估计方法,在小样本情境下评估诊断分类模型(DCMs)的性能。结果表明,贝叶斯估计在简约型DCMs中略微提升了被试分类的准确性,并且在小样本条件下相比ML方法展现出更稳定的项目参数恢复能力,尽管两种方法在参数不稳定性和边界问题方面仍面临挑战。
Fueled by the call for formative assessments, diagnostic classification models (DCMs) have recently gained popularity in psychometrics. Despite their potential for providing diagnostic information that aids in classroom instruction and students' learning, empirical applications of DCMs to classroom assessments have been highly limited. This is partly because how DCMs with different estimation methods perform in small sample contexts is not yet well-explored. Hence, this study aims to investigate the performance of respondent classification and item parameter estimation with a comprehensive simulation design that resembles classroom assessments using different estimation methods. The key findings are the following: (1) although the marked difference in respondent classification accuracy was not observed among the maximum likelihood (ML), Bayesian, and nonparametric methods, the Bayesian method provided slightly more accurate respondent classification in parsimonious DCMs than the ML method, and in complex DCMs, the ML method yielded the slightly better result than the Bayesian method; (2) while item parameter recovery was poor in both Bayesian and ML methods, the Bayesian method exhibited unstable slip values owing to the multimodality of their posteriors under complex DCMs, and the ML method produced irregular estimates that appear to be well-estimated due to a boundary problem under parsimonious DCMs.
研究动机与目标
- 评估诊断分类模型(DCMs)在典型课堂评估中常见的小样本情境下的表现。
- 比较最大似然(ML)、贝叶斯和非参数估计方法在被试分类准确性和项目参数估计方面的表现。
- 探究模型复杂度(简约型与复杂型DCMs)、样本量(20–40)以及项目数量对估计性能的影响。
- 为在数据有限的真实课堂环境中应用DCMs提供实用建议。
- 考察不同估计方法和模型类型下项目参数(失误率和猜测率)的稳定性和偏差。
提出的方法
- 采用五种DCMs(DINA、DINO、RRUM、CRUM和LCDM)进行综合性模拟研究,每种模型均基于其对应的真值模型生成。
- 在不同条件下生成模拟数据:样本量为20和40,项目数为10–30,属性数为3–5。
- 使用三种方法进行估计:最大似然(ML)、基于后验期望(EAP)的贝叶斯估计,以及非参数估计。
- 通过分类准确性、均方根误差(RMSE)以及项目参数估计的偏差(失误率和猜测率)来评估模型拟合度和参数恢复情况。
- 所有模拟中均正确指定了Q-matrix,以隔离估计方法和样本量的影响。
- 研究重点关注被试分类准确性与项目参数恢复,尤其关注小样本下的行为表现。
实验结果
研究问题
- RQ1在小样本量下,ML、贝叶斯和非参数估计方法在将被试分类到属性掌握模式方面表现如何比较?
- RQ2样本量(20–40)如何影响DCMs中被试分类准确性和项目参数估计的准确性?
- RQ3在小样本条件下,简约型与复杂型DCMs之间的性能差异是什么?
- RQ4在小样本中,ML与贝叶斯估计下估计的失误率和猜测率参数的稳定性和偏差如何?
- RQ5ML估计在何种情况下会出现边界问题?贝叶斯估计如何处理项目参数的多峰后验分布?
主要发现
- 在简约型DCMs中,贝叶斯估计相比ML提供了略高的被试分类准确性;而在复杂DCMs中,ML表现略优。
- 总体而言,项目参数恢复表现较差,ML与贝叶斯方法在小样本(N=20, 40)下对失误率和猜测率参数的RMSE均较高。
- ML估计出现了边界问题,在DINA和DINO模型中,失误率和猜测率参数收敛至0.0001,造成拟合良好的假象。
- 在复杂DCMs中,贝叶斯估计在CRUM和LCDM模型中对失误率参数产生了多峰后验分布,表明参数恢复存在不稳定性。
- 增加项目数量可降低ML与贝叶斯方法的偏差和RMSE,表明更长的测验能提升估计精度。
- 尽管存在局限性,贝叶斯估计仍推荐用于小样本课堂应用,因其分类结果更稳定,并且能够整合先验知识。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。