Skip to main content
QUICK REVIEW

[論文レビュー] Putting GPT-3's Creativity to the (Alternative Uses) Test

Claire E. Stevenson, Iris Smal|arXiv (Cornell University)|Jun 10, 2022
Creativity in Education and Neuroscience被引用数 54
ひとこと要約

この論文はGPT-3をGuilford's Alternative Uses Testで評価し、その独創性、有用性、驚き、柔軟性を人間の回答と比較し、人間は一般的にGPT-3より優れているが、GPT-3が追いつく可能性を示唆している。

ABSTRACT

AI large language models have (co-)produced amazing written works from newspaper articles to novels and poetry. These works meet the standards of the standard definition of creativity: being original and useful, and sometimes even the additional element of surprise. But can a large language model designed to predict the next text fragment provide creative, out-of-the-box, responses that still solve the problem at hand? We put Open AI's generative natural language model, GPT-3, to the test. Can it provide creative solutions to one of the most commonly used tests in creativity research? We assessed GPT-3's creativity on Guilford's Alternative Uses Test and compared its performance to previously collected human responses on expert ratings of originality, usefulness and surprise of responses, flexibility of each set of ideas as well as an automated method to measure creativity based on the semantic distance between a response and the AUT object in question. Our results show that -- on the whole -- humans currently outperform GPT-3 when it comes to creative output. But, we believe it is only a matter of time before GPT-3 catches up on this particular task. We discuss what this work reveals about human and AI creativity, creativity testing and our definition of creativity.

研究の動機と目的

  • GuilfordのAlternative Uses Test (AUT) に対して、創造的で独創的かつ有用な回答を生成するGPT-3の能力を評価する。
  • 独創性、有用性、驚き、柔軟性の面でGPT-3の出力を専門家評価の人間の回答と比較する。
  • AUTの回答に対する創造性指標として自動的な意味距離測度を検討する。
  • 人間とAIの創造性、創造性テスト、創造性の定義に関する含意を論じる。

提案手法

  • GuilfordのAlternative Uses Testの項目に対してGPT-3を適用して回答を生成する。
  • GPT-3の出力を、独創性、有用性、驚きについて専門家評価を行った人間の既存の回答と比較する。
  • 項目間での回答セットの柔軟性(GPT-3対人間)を評価する。
  • 回答とAUT対象物との意味距離に基づく自動創造性指標を評価する。
  • AIと人間の創造性の解釈、創造性の定義についての討論を提供する。

実験結果

リサーチクエスチョン

  • RQ1GPT-3のAUT回答は、独創性、有用性、驚きの面で専門家評価済みの人間の回答とどう比較されるか?
  • RQ2GPT-3の回答はAUT課題において人間の回答と同等の柔軟性を示すか?
  • RQ3意味距離ベースの自動創造性指標は、人間の評価と比較してAUT回答に対してどれほど有効か?
  • RQ4これらの結果はAIの創造性と現在の創造性テストのパラダイムについて何を示唆するか?

主な発見

  • 専門家の評価によれば、創造的な出力では人間がGPT-3を一般的に上回る。
  • GPT-3はある程度の創造的潜在性を示すが、独創性、有用性、驚きの面で人間の回答に遅れを取る。
  • この課題で今後ギャップを埋める可能性がGPT-3にはある。
  • 本研究は人間とAIの創造性および創造性テストの解釈に関する議論に情報を提供する。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。