[Paper Review] Artificial muses: Generative Artificial Intelligence Chatbots Have Risen to Human-Level Creativity
This study evaluates whether generative AI chatbots (GAI) achieve human-level creativity using the Alternative Uses Test (AUT), comparing 100 humans and six GAI models (including GPT-4) on idea fluency and originality. Results show no significant qualitative difference in creativity between top-tier GAI and humans, with 9.4% of humans exceeding GPT-4’s originality, suggesting GAI are strong creative assistants despite limitations in true intentionality.
A widespread view is that Artificial Intelligence cannot be creative. We tested this assumption by comparing human-generated ideas with those generated by six Generative Artificial Intelligence (GAI) chatbots: $alpa.\!ai$, $Copy.\!ai$, ChatGPT (versions 3 and 4), $Studio.\!ai$, and YouChat. Humans and a specifically trained AI independently assessed the quality and quantity of ideas. We found no qualitative difference between AI and human-generated creativity, although there are differences in how ideas are generated. Interestingly, 9.4 percent of humans were more creative than the most creative GAI, GPT-4. Our findings suggest that GAIs are valuable assistants in the creative process. Continued research and development of GAI in creative tasks is crucial to fully understand this technology's potential benefits and drawbacks in shaping the future of creativity. Finally, we discuss the question of whether GAIs are capable of being truly creative.
Motivation & Objective
- To test the widely held belief that AI cannot be creative by comparing GAI-generated ideas with human-generated ideas.
- To assess whether generative AI chatbots can produce ideas of comparable quality and originality to humans in everyday creative tasks.
- To investigate whether GAI can serve as effective assistants in the creative process, particularly in idea generation.
- To examine the extent to which GAI outputs meet the criteria of novelty and usefulness in a social context.
Proposed method
- Administered the Alternative Uses Test (AUT) to 100 human participants, asking them to generate multiple original uses for five common objects (e.g., tire, toothbrush).
- Collected responses from six GAI models: alpa.ai, Copy.ai, ChatGPT-3, ChatGPT-4, Studio.ai, and YouChat.
- Used both human raters and a separately trained AI model to assess the originality and fluency of all responses.
- Applied standardized scoring methods for divergent thinking, focusing on idea fluency (number of ideas) and originality (uniqueness relative to a normative sample).
- Conducted inter-rater reliability checks using the intraclass correlation coefficient (ICC) to ensure consistency in human and AI evaluations.
- Analyzed results using mixed-effects models to account for variability across objects and raters, with significance assessed at p < 0.05.
Experimental results
Research questions
- RQ1Do generative AI chatbots produce ideas that are qualitatively as creative as those generated by humans?
- RQ2Is there a measurable difference in originality and fluency between the most advanced GAI (GPT-4) and human participants?
- RQ3Can GAI be considered reliable assistants in the creative process, particularly in idea generation?
- RQ4To what extent do human participants surpass the most creative GAI output in terms of originality?
Key findings
- There was no significant qualitative difference in the originality or fluency of ideas generated by the most advanced GAI (GPT-4) and human participants.
- GPT-4 produced ideas that were statistically indistinguishable from human-generated ideas in terms of originality and quantity.
- 9.4% of human participants generated ideas more original than the most creative output from GPT-4, indicating that human creativity still exceeds GAI in a minority of cases.
- The study found that GAI outputs were consistently rated as novel and useful by both human and AI raters, supporting the perception of 'creative' output.
- The results suggest that GAI, particularly GPT-4, can serve as effective assistants in idea generation, especially in tasks requiring rapid ideation.
- Despite high performance, GAI outputs remain dependent on prompt quality and lack the full cognitive and emotional context of human creativity.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.