[论文解读] Leveraging LLM-Respondents for Item Evaluation: a Psychometric Analysis
本研究评估了大型语言模型(LLMs)作为教育评估中项目校准的合成被试的适用性。基于项目反应理论(IRT),研究发现GPT-3.5和LLM集成模型的响应模式与人类能力分布高度一致,且其项目参数与人类校准数据的相关性极高(r > 0.87);通过重采样进行数据增强后,相关性进一步提升至0.93,从而实现了可扩展、低成本的项目评估,适用于形成性与总结性评估。
Effective educational measurement relies heavily on the curation of well-designed item pools (i.e., possessing the right psychometric properties). However, item calibration is time-consuming and costly, requiring a sufficient number of respondents for the response process. We explore using six different LLMs (GPT-3.5, GPT-4, Llama 2, Llama 3, Gemini-Pro, and Cohere Command R Plus) and various combinations of them using sampling methods to produce responses with psychometric properties similar to human answers. Results show that some LLMs have comparable or higher proficiency in College Algebra than college students. No single LLM mimics human respondents due to narrow proficiency distributions, but an ensemble of LLMs can better resemble college students' ability distribution. The item parameters calibrated by LLM-Respondents have high correlations (e.g. > 0.8 for GPT-3.5) compared to their human calibrated counterparts, and closely resemble the parameters of the human subset (e.g. 0.02 Spearman correlation difference). Several augmentation strategies are evaluated for their relative performance, with resampling methods proving most effective, enhancing the Spearman correlation from 0.89 (human only) to 0.93 (augmented human).
研究动机与目标
- 评估LLMs是否能够通过心理测量建模在教育评估中模拟人类响应模式。
- 比较使用LLM生成响应与人类响应校准的项目心理测量特性。
- 评估结合LLM被试与人类数据的数据增强策略,以提升校准精度。
- 探索LLMs在形成性与总结性评估中实现快速、低成本项目筛选的可行性。
提出的方法
- 六种LLM(GPT-3.5、GPT-4、Llama 2、Llama 3、Gemini-Pro、Cohere Command R Plus)被用于处理OpenStax提供的20道大学代数题目,每种模型生成150条响应。
- 应用项目反应理论(IRT)对人类和LLM生成响应数据中的项目参数(难度、区分度)进行校准。
- 计算LLM校准与人类校准项目参数之间的Spearman与Pearson相关系数,以评估相似性。
- 测试了三种增强策略:直接组合、LLM数据重采样,以及重采样LLM数据与人类数据的1:1混合。
- 通过重采样优化LLM被试的能力分布,降低偏态,更好地匹配人类的双峰分布模式。
- 通过相关性指标、均方根误差(RMSE)以及与人类基线分布的比较,评估模型性能。

实验结果
研究问题
- RQ1RQ1:哪种LLM或LLM组合在IRT测量的数学能力上最能模拟人类被试的行为?
- RQ2RQ2:使用LLM被试校准的项目心理测量特性与使用人类被试校准的项目相比如何?
- RQ3RQ3:能否通过将LLM被试数据与人类响应数据结合,获得与更大规模人类样本相当的项目参数?
主要发现
- GPT-3.5在与人类被试的相似性方面表现最佳,与人类校准项目参数的Spearman相关系数达到0.87。
- 采用重采样策略的LLM集成模型将LLM与人类校准参数之间的相关性提升至0.93,高于仅使用人类数据时的0.89。
- 单一LLM无法完全复制人类能力分布,因其能力分布范围过窄,但结合重采样策略的LLM集成模型更接近人类的双峰分布。
- 最佳增强策略——重采样LLM与人类数据的1:1混合——使Spearman相关系数提高0.04,Pearson相关系数提高0.01,尽管也导致RMSE上升。
- LLM被试(GPT-3.5生成的150条响应)生成的项目参数与人类数据的相关性达到0.87,与50名人类被试样本的性能相当。
- 本研究首次通过IRT揭示了多种LLM之间的全新能力分布,突破了仅依赖点估计的局限,实现了对完整能力分布的建模。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。