Skip to main content
QUICK REVIEW

[Paper Review] Artificial Intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry

Nils Köbis, Luca Mossink|arXiv (Cornell University)|May 20, 2020
Topic Modeling4 citations
TL;DR

This study investigates whether people can distinguish AI-generated poetry from human-written poetry using GPT-2. In a controlled experiment with 830 participants, people failed to reliably detect AI-generated poems when the best output was selected (Human-in-the-loop), but could do so when random outputs were used (Human-out-of-the-loop), indicating advanced human-like text generation. Participants also showed a slight aversion to AI poetry regardless of transparency about its origin.

ABSTRACT

The release of openly available, robust natural language generation algorithms (NLG) has spurred much public attention and debate. One reason lies in the algorithms' purported ability to generate human-like text across various domains. Empirical evidence using incentivized tasks to assess whether people (a) can distinguish and (b) prefer algorithm-generated versus human-written text is lacking. We conducted two experiments assessing behavioral reactions to the state-of-the-art Natural Language Generation algorithm GPT-2 (Ntotal = 830). Using the identical starting lines of human poems, GPT-2 produced samples of poems. From these samples, either a random poem was chosen (Human-out-of-the-loop) or the best one was selected (Human-in-the-loop) and in turn matched with a human-written poem. In a new incentivized version of the Turing Test, participants failed to reliably detect the algorithmically-generated poems in the Human-in-the-loop treatment, yet succeeded in the Human-out-of-the-loop treatment. Further, people reveal a slight aversion to algorithm-generated poetry, independent on whether participants were informed about the algorithmic origin of the poem (Transparency) or not (Opacity). We discuss what these results convey about the performance of NLG algorithms to produce human-like text and propose methodologies to study such learning algorithms in human-agent experimental settings.

Motivation & Objective

  • To assess whether people can reliably distinguish AI-generated poetry from human-written poetry.
  • To examine preferences for AI-generated versus human-written poetry in controlled experimental settings.
  • To evaluate the impact of transparency (knowing the source) on perception and preference of AI-generated text.
  • To test the effectiveness of state-of-the-art NLG models like GPT-2 in producing human-like poetic text.
  • To develop methodologies for studying human-agent interactions in experimental settings involving generative AI.

Proposed method

  • Conducted two incentivized experiments with 830 participants using identical starting lines from human poems.
  • Used GPT-2 to generate poetry samples from the same starting lines, selecting either the best output (Human-in-the-loop) or a random output (Human-out-of-the-loop).
  • Matched each AI-generated poem with a human-written poem of similar quality for comparison.
  • Implemented a novel incentivized Turing Test variant to assess detection accuracy and preference.
  • Varied conditions between Transparency (participants informed of AI origin) and Opacity (no information given) to assess bias.
  • Collected behavioral responses on detection accuracy and preference using controlled, double-blind experimental design.

Experimental results

Research questions

  • RQ1Can participants reliably detect AI-generated poetry when the best output from GPT-2 is used?
  • RQ2Does the selection method (best vs. random output) affect participants' ability to distinguish AI from human poetry?
  • RQ3Do participants show a preference for human-written or AI-generated poetry, and does this preference depend on transparency about the source?
  • RQ4How does transparency about the algorithmic origin of poetry influence perception and evaluation of poetic quality?
  • RQ5To what extent do state-of-the-art NLG models like GPT-2 produce text that is indistinguishable from human writing in creative domains like poetry?

Key findings

  • Participants failed to reliably detect AI-generated poetry in the Human-in-the-loop condition, where the best GPT-2 output was selected, indicating strong human-likeness.
  • Participants successfully detected AI-generated poetry in the Human-out-of-the-loop condition, where random GPT-2 outputs were used, suggesting variability in output quality affects detectability.
  • A slight but consistent aversion to AI-generated poetry was observed, regardless of whether participants were informed of the algorithmic origin (Transparency) or not (Opacity).
  • The study provides empirical evidence that state-of-the-art NLG models like GPT-2 can produce poetry that is nearly indistinguishable from human writing under optimal conditions.
  • The results demonstrate that current NLG systems can produce text that passes basic human perception tests in creative domains, challenging assumptions about human uniqueness in artistic expression.
  • The study highlights the need for new methodological frameworks to study human-agent interactions involving generative AI in experimental settings.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.