Skip to main content
QUICK REVIEW

[Paper Review] Large Language Models show both individual and collective creativity comparable to humans

Luning Sun, Yuzhuo Yuan|arXiv (Cornell University)|Dec 4, 2024
Machine Learning in Materials ScienceMaterials Science3 citations
TL;DR

This study evaluates the creative capabilities of Large Language Models (LLMs) across 13 tasks spanning ideation, problem solving, and creative writing, comparing both individual and collective performance against humans. It finds that top LLMs (Claude and GPT-4) achieve human-level creativity in divergent thinking and problem solving, with collective LLM outputs matching 8–10 human collaborators, suggesting LLMs can rival small human teams in future work settings.

ABSTRACT

Artificial intelligence has, so far, largely automated routine tasks, but what does it mean for the future of work if Large Language Models (LLMs) show creativity comparable to humans? To measure the creativity of LLMs holistically, the current study uses 13 creative tasks spanning three domains. We benchmark the LLMs against individual humans, and also take a novel approach by comparing them to the collective creativity of groups of humans. We find that the best LLMs (Claude and GPT-4) rank in the 52nd percentile against humans, and overall LLMs excel in divergent thinking and problem solving but lag in creative writing. When questioned 10 times, an LLM's collective creativity is equivalent to 8-10 humans. When more responses are requested, two additional responses of LLMs equal one extra human. Ultimately, LLMs, when optimally applied, may compete with a small group of humans in the future of work.

Motivation & Objective

  • To assess whether Large Language Models (LLMs) exhibit creativity comparable to humans across diverse domains.
  • To evaluate LLMs not only individually but also collectively, simulating team-like performance.
  • To benchmark LLMs against both individual humans and groups of humans on a broad set of creative tasks.
  • To understand the scalability of LLM creativity when multiple responses are aggregated.
  • To explore the implications of LLMs for the future of work, particularly in creative and collaborative roles.

Proposed method

  • The study employs 13 creative tasks across three domains: ideation, problem solving, and creative writing.
  • LLMs are evaluated individually and in ensembles—by generating 10 responses per task to simulate collective output.
  • Human performance is measured through individual and group responses, with groups consisting of 8–10 participants.
  • Creativity is assessed using standardized metrics across tasks, including originality, diversity, and feasibility.
  • Statistical benchmarks compare LLM performance to individual humans and human groups, using percentiles to rank performance.
  • The study uses GPT-4 and Claude as leading LLMs, with results analyzed for consistency and scalability.

Experimental results

Research questions

  • RQ1Can LLMs achieve creativity levels comparable to individual humans across diverse creative tasks?
  • RQ2How does the collective creativity of LLMs, when generating multiple responses, compare to that of human groups?
  • RQ3In which creative domains do LLMs outperform or underperform relative to humans?
  • RQ4To what extent can aggregated LLM outputs substitute for human team collaboration in creative work?
  • RQ5What is the scalability of LLM creativity as the number of generated responses increases?

Key findings

  • The best LLMs (Claude and GPT-4) rank in the 52nd percentile compared to individual humans, indicating near-human-level creativity.
  • LLMs excel in divergent thinking and problem-solving tasks, outperforming humans in generating diverse and feasible solutions.
  • LLMs underperform in creative writing tasks, where human originality and narrative depth remain superior.
  • When queried 10 times, an LLM’s collective output matches the creativity of 8–10 human collaborators.
  • For every two additional LLM responses, the equivalent of one extra human contributor is added to the collective output.
  • LLMs, when optimally applied, may compete with small human teams in future creative and collaborative work environments.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.