[论文解读] Automated Reading Passage Generation with OpenAI's Large Language Model
本文提出了一种使用OpenAI的GPT-3语言模型进行自动化阅读段落生成的方法,通过精心设计的提示(prompts)生成与原文在内容、结构和Lexile分数上一致的四年级水平段落。该方法结合了AI生成与人工修订及评估,生成的文本在连贯性、可读性和年龄适宜性方面表现优异。
The widespread usage of computer-based assessments and individualized learning platforms has resulted in an increased demand for the rapid production of high-quality items. Automated item generation (AIG), the process of using item models to generate new items with the help of computer technology, was proposed to reduce reliance on human subject experts at each step of the process. AIG has been used in test development for some time. Still, the use of machine learning algorithms has introduced the potential to improve the efficiency and effectiveness of the process greatly. The approach presented in this paper utilizes OpenAI's latest transformer-based language model, GPT-3, to generate reading passages. Existing reading passages were used in carefully engineered prompts to ensure the AI-generated text has similar content and structure to a fourth-grade reading passage. For each prompt, we generated multiple passages, the final passage was selected according to the Lexile score agreement with the original passage. In the final round, the selected passage went through a simple revision by a human editor to ensure the text was free of any grammatical and factual errors. All AI-generated passages, along with original passages were evaluated by human judges according to their coherence, appropriateness to fourth graders, and readability.
研究动机与目标
- 为计算机化评估和个性化学习平台中快速增长的快速、高质量阅读段落需求提供解决方案。
- 通过使用大语言模型自动化生成过程,减少在阅读段落开发中对人类专家的依赖。
- 确保生成的段落与真实四年级阅读材料在语言复杂度和结构上保持一致。
- 通过人工判断评估生成段落的连贯性、可读性和年龄适宜性。
- 建立一个可扩展的流水线,结合大语言模型生成与最少的人工监督,实现教育场景中的实际应用。
提出的方法
- 利用OpenAI的GPT-3——一种基于Transformer的语言模型——从精心设计的提示中生成阅读段落。
- 基于现有的四年级阅读段落设计提示,引导模型匹配内容和结构模式。
- 每个提示生成多个段落,并根据生成段落与原文的Lexile分数一致性选择最终版本。
- 通过人工编辑审查修正所选段落中的语法和事实性错误。
- 由人工评委对所有生成段落和原始段落进行连贯性、可读性和年龄适宜性的评估。
- 通过与典型四年级课程标准对齐,确保内容在教育标准上的一致性。
实验结果
研究问题
- RQ1GPT-3能否生成在连贯性和结构上与真实四年级段落相似的阅读段落?
- RQ2AI生成的段落与原始段落在Lexile水平上有多大的匹配度?
- RQ3与人工编写的原始段落相比,人工评委如何评价AI生成段落的可读性和年龄适宜性?
- RQ4人工修订在提升AI生成阅读段落质量方面起到了什么作用?
- RQ5基于提示的微调方法结合GPT-3能否在大规模上持续生成高质量的教育类段落?
主要发现
- AI生成的段落在Lexile分数上与原始段落高度一致,表明文本复杂度相当。
- 人工评委评定生成的段落具有高度连贯性,并且适合四年级学生阅读。
- 可读性评估证实,生成的段落符合目标年级的语言期望。
- 人工修订显著提升了最终输出在语法准确性和事实正确性方面的表现。
- 提示工程与最少的人工监督相结合,生成的段落在关键评估维度上与人工编写的段落质量无异。
- 该方法在大规模自动化生成教育类阅读材料方面展现出可扩展性和实际可行性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。