[论文解读] Playing with words: Comparing the vocabulary and lexical diversity of ChatGPT and humans
本研究通过RTTR和Mass等指标,比较了ChatGPT与人类在多个任务中生成文本的词汇量和词汇丰富度。研究发现,ChatGPT始终使用的独特词汇更少,词汇多样性更低,表明若AI生成内容主导话语,可能对语言演化和词汇使用造成长期影响。
The introduction of Artificial Intelligence (AI) generative language models such as GPT (Generative Pre-trained Transformer) and tools such as ChatGPT has triggered a revolution that can transform how text is generated. This has many implications, for example, as AI-generated text becomes a significant fraction of the text, would this have an effect on the language capabilities of readers and also on the training of newer AI tools? Would it affect the evolution of languages? Focusing on one specific aspect of the language: words; will the use of tools such as ChatGPT increase or reduce the vocabulary used or the lexical richness? This has implications for words, as those not included in AI-generated content will tend to be less and less popular and may eventually be lost. In this work, we perform an initial comparison of the vocabulary and lexical richness of ChatGPT and humans when performing the same tasks. In more detail, two datasets containing the answers to different types of questions answered by ChatGPT and humans, and a third dataset in which ChatGPT paraphrases sentences and questions are used. The analysis shows that ChatGPT tends to use fewer distinct words and lower lexical richness than humans. These results are very preliminary and additional datasets and ChatGPT configurations have to be evaluated to extract more general conclusions. Therefore, further research is needed to understand how the use of ChatGPT and more broadly generative AI tools will affect the vocabulary and lexical richness in different types of text and languages.
研究动机与目标
- 调查在执行相同语言任务时,ChatGPT的词汇量和词汇丰富度与人类的对比情况。
- 评估AI生成文本是否可能影响语言的未来发展,特别是词汇频率和词汇多样性方面。
- 识别由于AI生成内容中词汇丰富度降低而可能导致的词汇萎缩风险。
- 为未来研究生成式AI工具对人类语言使用和发展所产生的语言影响奠定基础。
- 倡导建立标准化数据集和自动化工具,以大规模评估AI的词汇表现。
提出的方法
- 本研究分析了三个数据集:人类和ChatGPT对不同类型问题的回答,以及ChatGPT对人类问题的改写版本。
- 通过RTTR(根型词比)和Mass指标量化词汇丰富度,这两个指标衡量的是独特词汇与总词汇数的比率。
- 在二次分析中去除停用词,以隔离内容词(名词、动词、形容词、副词),并评估核心词汇多样性。
- 在不同任务中对比人类与ChatGPT的输出,通过可视化RTTR和Mass得分来突出差异。
- 分析聚焦于独特词汇数量和词汇多样性趋势,使用统计指标比较各数据集间的性能表现。
- 作者提出未来可通过API实现自动化,以在书籍、提示、模型版本和语言之间大规模扩展测试。
实验结果
研究问题
- RQ1在不同问题类型下,ChatGPT使用的独特词汇数量与人类相比如何?
- RQ2在RTTR和Mass指标下,ChatGPT的词汇丰富度在多大程度上低于人类?
- RQ3停用词与内容词如何影响ChatGPT与人类文本之间词汇多样性的比较?
- RQ4AI生成文本中词汇多样性降低对语言学习和语言未来演化有何影响?
- RQ5随着AI生成内容的使用日益增加,这将如何影响词汇频率以及不常用或已过时词汇的存续?
主要发现
- 在所有测试数据集中,ChatGPT使用的独特词汇显著少于人类,表明其词汇范围更窄。
- RTTR指标显示,ChatGPT的词汇丰富度始终低于人类,其值在不同任务中范围为0.0112至0.0130。
- 去除停用词后,ChatGPT的词汇丰富度仍较低,RTTR值在0.0129至0.0130之间,证实其内容词的多样性减少。
- Mass指标进一步证实,ChatGPT生成的文本词汇丰富度较低,得分越低表明词汇多样性越弱。
- 结果在不同任务中保持一致,包括问答和改写任务,表明ChatGPT在词汇范围上存在系统性局限。
- 本研究得出结论,未来需进一步研究这些发现是否可推广至其他AI模型、语言和配置。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。