Skip to main content
QUICK REVIEW

[论文解读] Artificial Intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry

Nils Köbis, Luca Mossink|arXiv (Cornell University)|May 20, 2020
Topic Modeling被引用 4
一句话总结

本研究调查了人们是否能通过使用 GPT-2 区分人工智能生成的诗歌与人类创作的诗歌。在包含 830 名参与者的受控实验中,当选择最佳输出(人机协同)时,人们无法可靠检测出人工智能生成的诗歌;但当使用随机输出(无人参与)时,人们则能够做到,表明文本生成已达到高度类人的水平。此外,无论是否知晓其来源,参与者对人工智能诗歌均表现出轻微的排斥倾向。

ABSTRACT

The release of openly available, robust natural language generation algorithms (NLG) has spurred much public attention and debate. One reason lies in the algorithms' purported ability to generate human-like text across various domains. Empirical evidence using incentivized tasks to assess whether people (a) can distinguish and (b) prefer algorithm-generated versus human-written text is lacking. We conducted two experiments assessing behavioral reactions to the state-of-the-art Natural Language Generation algorithm GPT-2 (Ntotal = 830). Using the identical starting lines of human poems, GPT-2 produced samples of poems. From these samples, either a random poem was chosen (Human-out-of-the-loop) or the best one was selected (Human-in-the-loop) and in turn matched with a human-written poem. In a new incentivized version of the Turing Test, participants failed to reliably detect the algorithmically-generated poems in the Human-in-the-loop treatment, yet succeeded in the Human-out-of-the-loop treatment. Further, people reveal a slight aversion to algorithm-generated poetry, independent on whether participants were informed about the algorithmic origin of the poem (Transparency) or not (Opacity). We discuss what these results convey about the performance of NLG algorithms to produce human-like text and propose methodologies to study such learning algorithms in human-agent experimental settings.

研究动机与目标

  • 评估人们是否能可靠区分人工智能生成的诗歌与人类创作的诗歌。
  • 在受控实验环境中,考察人们对人工智能生成诗歌与人类创作诗歌的偏好。
  • 评估透明度(知晓来源)对人工智能生成文本的感知与偏好影响。
  • 测试最先进的自然语言生成模型(如 GPT-2)在生成类人化诗歌文本方面的有效性。
  • 开发用于在涉及生成式人工智能的实验环境中研究人机交互的方法论。

提出的方法

  • 开展两项有激励的实验,共招募 830 名参与者,使用来自人类诗歌的相同起始句。
  • 使用 GPT-2 从相同起始句生成诗歌样本,选择最佳输出(人机协同)或随机输出(无人参与)。
  • 将每首人工智能生成的诗歌与质量相近的人类创作诗歌配对进行对比。
  • 采用一种新型有激励的图灵测试变体,以评估检测准确率与偏好。
  • 在透明度(参与者被告知来源为人工智能)与不透明度(未提供信息)之间设置不同条件,以评估偏见影响。
  • 通过受控的双盲实验设计,收集关于检测准确率与偏好的行为反应数据。

实验结果

研究问题

  • RQ1当使用 GPT-2 的最佳输出时,参与者能否可靠检测出人工智能生成的诗歌?
  • RQ2选择方法(最佳输出 vs. 随机输出)是否会影响参与者区分人工智能与人类诗歌的能力?
  • RQ3参与者是否对人类创作或人工智能生成的诗歌有偏好,且该偏好是否取决于是否知晓来源?
  • RQ4对诗歌算法来源的透明度如何影响其感知与质量评价?
  • RQ5在诗歌等创意领域,最先进的自然语言生成模型(如 GPT-2)在多大程度上能生成与人类写作几乎无法区分的文本?

主要发现

  • 在人机协同条件下(即选择 GPT-2 最佳输出时),参与者无法可靠检测出人工智能生成的诗歌,表明其具有极强的类人特征。
  • 在无人参与条件下(即使用随机 GPT-2 输出时),参与者成功检测出人工智能生成的诗歌,表明输出质量的差异会影响可检测性。
  • 无论参与者是否知晓来源(透明度或不透明度),均表现出对人工智能诗歌的轻微但一致的排斥倾向。
  • 本研究提供了实证证据,表明最先进的自然语言生成模型(如 GPT-2)在最优条件下可生成几乎无法与人类写作区分的诗歌。
  • 结果表明,当前的自然语言生成系统可在创意领域通过基本的人类感知测试,对艺术表达中人类独特性的假设构成挑战。
  • 本研究强调了在涉及生成式人工智能的实验环境中,亟需建立新的方法论框架以研究人机交互。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。