Skip to main content
QUICK REVIEW

[论文解读] Putting GPT-3's Creativity to the (Alternative Uses) Test

Claire E. Stevenson, Iris Smal|arXiv (Cornell University)|Jun 10, 2022
Creativity in Education and Neuroscience被引用 54
一句话总结

本论文评估 GPT-3 在 Guilford's Alternative Uses Test 中的创造力,并将其原创性、实用性、惊奇度和灵活性与人类回答进行比较,结果是人类通常优于 GPT-3,但建议 GPT-3 可能会赶上。

ABSTRACT

AI large language models have (co-)produced amazing written works from newspaper articles to novels and poetry. These works meet the standards of the standard definition of creativity: being original and useful, and sometimes even the additional element of surprise. But can a large language model designed to predict the next text fragment provide creative, out-of-the-box, responses that still solve the problem at hand? We put Open AI's generative natural language model, GPT-3, to the test. Can it provide creative solutions to one of the most commonly used tests in creativity research? We assessed GPT-3's creativity on Guilford's Alternative Uses Test and compared its performance to previously collected human responses on expert ratings of originality, usefulness and surprise of responses, flexibility of each set of ideas as well as an automated method to measure creativity based on the semantic distance between a response and the AUT object in question. Our results show that -- on the whole -- humans currently outperform GPT-3 when it comes to creative output. But, we believe it is only a matter of time before GPT-3 catches up on this particular task. We discuss what this work reveals about human and AI creativity, creativity testing and our definition of creativity.

研究动机与目标

  • 评估 GPT-3 在 Guilford 的替代用途测试(AUT)中生成具有创造性、原创性和实用性的回答的能力。
  • 将 GPT-3 的输出与专家评分的人类回答在原创性、实用性和惊奇度以及灵活性方面进行比较。
  • 研究一种基于语义距离的自动化创造力度量在 AUT 回应中的应用效果。
  • 讨论 AI 与人类创造力、创造力测试以及创造力定义的含义。

提出的方法

  • 将 GPT-3 应用于 Guilford 的替代用途测试中的条目生成回答。
  • 将 GPT-3 的输出与先前收集的人类回答进行比较,专家对原创性、实用性和惊奇度进行评分。
  • 评估在条目间的回答集的灵活性(GPT-3 与人类)。
  • 基于回答与 AUT 对象之间的语义距离的自动创造力度量进行评估。
  • 就 AI 与人类创造力的解释以及创造力的定义提供讨论。

实验结果

研究问题

  • RQ1GPT-3 的 AUT 回答在原创性、实用性和惊奇度方面与专家评分的人类回答相比如何?
  • RQ2GPT-3 的回答在 AUT 任务中是否表现出与人类回答相当的灵活性?
  • RQ3基于语义距离的自动创造力度量在 AUT 回应中的有效性,与人类评分相比如何?
  • RQ4这些结果对 AI 创造力与当前创造力测试范式有何启示?

主要发现

  • 根据专家评分,人类在创造性输出方面通常优于 GPT-3。
  • GPT-3 展现出一定的创造潜力,但在原创性、实用性和惊奇度方面落后于人类回答。
  • 未来有可能缩小这一任务中的差距。
  • 这项工作为关于人类与 AI 创造力以及创造力测试的解释提供了参考。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。