Skip to main content
QUICK REVIEW

[Paper Review] Divergent Creativity in Humans and Large Language Models

Antoine Bellemare-Pepin, François Lespinasse|arXiv (Cornell University)|May 13, 2024
Creativity in Education and Neuroscience15 citations
TL;DR

The paper systematically compares semantic diversity in state-of-the-art LLMs to a large human dataset, showing LLMs can exceed average humans on divergent tasks but do not surpass highly creative individuals, and it proposes benchmarking and methods to improve semantic diversity.

ABSTRACT

The recent surge of Large Language Models (LLMs) has led to claims that they are approaching a level of creativity akin to human capabilities. This idea has sparked a blend of excitement and apprehension. However, a critical piece that has been missing in this discourse is a systematic evaluation of LLMs' semantic diversity, particularly in comparison to human divergent thinking. To bridge this gap, we leverage recent advances in computational creativity to analyze semantic divergence in both state-of-the-art LLMs and a substantial dataset of 100,000 humans. We found evidence that LLMs can surpass average human performance on the Divergent Association Task, and approach human creative writing abilities, though they fall short of the typical performance of highly creative humans. Notably, even the top performing LLMs are still largely surpassed by highly creative individuals, underscoring a ceiling that current LLMs still fail to surpass. Our human-machine benchmarking framework addresses the polemic surrounding the imminent replacement of human creative labour by AI, disentangling the quality of the respective creative linguistic outputs using established objective measures. While prompting deeper exploration of the distinctive elements of human inventive thought compared to those of AI systems, we lay out a series of techniques to improve their outputs with respect to semantic diversity, such as prompt design and hyper-parameter tuning.

Motivation & Objective

  • Assess semantic diversity of state-of-the-art LLMs versus a large human dataset on divergent thinking tasks.
  • Quantify where LLMs stand relative to average and highly creative humans in divergent association and creative writing.
  • Provide a human–machine benchmarking framework to evaluate creative linguistic outputs with objective measures.
  • Offer techniques, such as prompt design and hyper-parameter tuning, to improve LLM semantic diversity.

Proposed method

  • Apply computational creativity methods to measure semantic divergence in LLM outputs.
  • Use Divergent Association Task and creative writing benchmarks to compare against 100,000 human data points.
  • Benchmark LLM performance across multiple prompting strategies and model configurations.
  • Analyze outputs with established objective measures of creativity and linguistic diversity.

Experimental results

Research questions

  • RQ1Do LLMs surpass average humans on divergent thinking tasks?
  • RQ2Do LLMs approach or exceed the creativity of highly creative humans?
  • RQ3What are the key differences in the qualitative and quantitative aspects of human and AI creative outputs?
  • RQ4What prompting and hyper-parameter strategies can increase LLM semantic diversity?

Key findings

  • LLMs can surpass average human performance on the Divergent Association Task.
  • LLMs approach human creative writing abilities but do not reach the typical performance of highly creative humans.
  • Even top-performing LLMs are largely surpassed by highly creative individuals, indicating a ceiling for current models.
  • A human–machine benchmarking framework helps disentangle output quality using objective measures.
  • The paper suggests techniques such as prompt design and hyper-parameter tuning to improve semantic diversity.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.