[论文解读] Divergent Creativity in Humans and Large Language Models
这篇论文系统地将最前沿大语言模型的语义多样性与大规模人类数据集进行比较,结果表明 LLMs 在发散任务上可以超越普通人,但尚未超过高度有创造力的个体;并提出基准测试与方法以提升语义多样性。
The recent surge of Large Language Models (LLMs) has led to claims that they are approaching a level of creativity akin to human capabilities. This idea has sparked a blend of excitement and apprehension. However, a critical piece that has been missing in this discourse is a systematic evaluation of LLMs' semantic diversity, particularly in comparison to human divergent thinking. To bridge this gap, we leverage recent advances in computational creativity to analyze semantic divergence in both state-of-the-art LLMs and a substantial dataset of 100,000 humans. We found evidence that LLMs can surpass average human performance on the Divergent Association Task, and approach human creative writing abilities, though they fall short of the typical performance of highly creative humans. Notably, even the top performing LLMs are still largely surpassed by highly creative individuals, underscoring a ceiling that current LLMs still fail to surpass. Our human-machine benchmarking framework addresses the polemic surrounding the imminent replacement of human creative labour by AI, disentangling the quality of the respective creative linguistic outputs using established objective measures. While prompting deeper exploration of the distinctive elements of human inventive thought compared to those of AI systems, we lay out a series of techniques to improve their outputs with respect to semantic diversity, such as prompt design and hyper-parameter tuning.
研究动机与目标
- 评估最前沿 LLMs 与大规模人类数据集在发散性思维任务中的语义多样性。
- 量化 LLMs 相对于普通与高度有创造力的人类在发散联想和创造性写作方面的地位。
- 提供一个人机基准框架,以使用客观测量评估创造性语言输出。
- 提供如提示词设计和超参数调优等技术,以提升 LLM 的语义多样性。
提出的方法
- 应用计算创造力方法来衡量 LLM 输出的语义发散性。
- 使用 Divergent Association Task 与创造性写作基准来与10万条人类数据进行比较。
- 在多种提示策略和模型配置下对 LLM 表现进行基准测试。
- 用已建立的创造性与语言多样性的客观指标分析输出结果。
实验结果
研究问题
- RQ1LLMs 是否在发散性思维任务上超过普通人?
- RQ2LLMs 是否接近或超过高度有创造力的人类的创造力?
- RQ3在人类与 AI 的创造性输出在定性和定量方面的关键差异是什么?
- RQ4哪些 prompting 与超参数策略可以提高 LLM 的语义多样性?
主要发现
- LLMs 在 Divergent Association Task 上可以超越普通人表现。
- LLMs 在创造性写作方面接近人类,但未达到高度有创造力的人类的典型水平。
- 即使是表现最出色的 LLM,也往往被高度有创造力的个体压制,表明当前模型存在天花板。
- 一个人机基准框架有助于用客观指标解析输出质量。
- 论文提出了如提示词设计和超参数调优等提升语义多样性的技术。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。