[论文解读] What has ChatGPT read? The origins of archaeological citations used by a generative artificial intelligence application
本研究通过填空分析法调查了ChatGPT生成的考古学引用的来源,以评估该模型在训练期间是否访问过真实的学术文献。研究发现,尽管ChatGPT提供的引用看似合理,但大多数为虚构;真实引用很可能源自维基百科而非原始文献,这引发了对AI生成学术内容中数据质量和可靠性的担忧。
The public release of ChatGPT has resulted in considerable publicity and has led to wide-spread discussion of the usefulness and capabilities of generative AI language models. Its ability to extract and summarise data from textual sources and present them as human-like contextual responses makes it an eminently suitable tool to answer questions users might ask. This paper tested what archaeological literature appears to have been included in ChatGPT's training phase. While ChatGPT offered seemingly pertinent references, a large percentage proved to be fictitious. Using cloze analysis to make inferences on the sources 'memorised' by a generative AI model, this paper was unable to prove that ChatGPT had access to the full texts of the genuine references. It can be shown that all references provided by ChatGPT that were found to be genuine have also been cited on Wikipedia pages. This strongly indicates that the source base for at least some of the data is found in those pages. The implications of this in relation to data quality are discussed.
研究动机与目标
- 确定ChatGPT所引用的考古学文献的来源。
- 评估ChatGPT在训练阶段是否访问过真实学术文献的全文版本。
- 评估AI生成引用在学术与文化遗产背景下的可靠性。
- 调查维基百科是否为ChatGPT在考古学与文化遗产领域训练数据的主要来源。
提出的方法
- 通过填空分析法推断ChatGPT基于其完成引用片段的能力,判断哪些来源被其‘记忆’。
- 收集并交叉核对ChatGPT生成的引用与已知的考古学及文化遗产管理领域的学术文献。
- 通过检查其是否存在于学术数据库(如JSTOR)、Google图书及Archive.org,验证引用的真实性。
- 将引用来源与维基百科页面的引用进行比对,以识别潜在的数据来源路径。
- 在多个主题和模型版本上重复查询,以测试响应的一致性与可靠性。
- 通过评估560个引用中作者、标题、年份及来源完整性的准确性,分析引用的准确性。
实验结果
研究问题
- RQ1ChatGPT生成的考古学引用中,事实准确与虚构的比例各占多少?
- RQ2哪些来源最有可能为ChatGPT在考古学引用方面的训练数据做出贡献?
- RQ3ChatGPT引用的真实文献是否出现在对应的维基百科页面上,表明其可能从维基百科而非原始文献中学习?
- RQ4ChatGPT在训练期间在多大程度上访问了所引用学术著作的全文版本?
- RQ5ChatGPT在重复查询和不同模型版本间,其引用的一致性如何?
主要发现
- 在分析的560个引用中,48.8%为虚构或捏造,表明AI生成引用存在较高的幻觉率。
- 仅有32.1%的引用在作者、标题和年份信息上完全准确。
- ChatGPT引用的所有真实文献均在对应维基百科页面上被引用,表明维基百科很可能是其训练数据来源。
- 未发现证据表明ChatGPT在训练期间访问过真实学术文献的全文版本。
- 在澳大利亚考古学领域,84.5%的引用为虚构,表明该领域存在特别高的错误率。
- 该模型在重复查询中表现出一致的引用错误,再生响应常产生不同但同样不准确的引用。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。