[논문 리뷰] Putting GPT-3's Creativity to the (Alternative Uses) Test
이 논문은 Guilford의 Alternative Uses Test(AUT)에서 GPT-3를 평가하고 창의성, 유용성, 놀라움, 유연성을 인간 응답과 비교하여 일반적으로 인간이 GPT-3를 능가하지만 GPT-3가 격차를 따라잡을 수 있음을 시사한다.
AI large language models have (co-)produced amazing written works from newspaper articles to novels and poetry. These works meet the standards of the standard definition of creativity: being original and useful, and sometimes even the additional element of surprise. But can a large language model designed to predict the next text fragment provide creative, out-of-the-box, responses that still solve the problem at hand? We put Open AI's generative natural language model, GPT-3, to the test. Can it provide creative solutions to one of the most commonly used tests in creativity research? We assessed GPT-3's creativity on Guilford's Alternative Uses Test and compared its performance to previously collected human responses on expert ratings of originality, usefulness and surprise of responses, flexibility of each set of ideas as well as an automated method to measure creativity based on the semantic distance between a response and the AUT object in question. Our results show that -- on the whole -- humans currently outperform GPT-3 when it comes to creative output. But, we believe it is only a matter of time before GPT-3 catches up on this particular task. We discuss what this work reveals about human and AI creativity, creativity testing and our definition of creativity.
연구 동기 및 목표
- Guilford의 Alternative Uses Test (AUT)에 대해 GPT-3가 창의적이고 독창적이며 유용한 응답을 생성할 수 있는 능력을 평가한다.
- GPT-3 출력물을 독창성, 유용성, 놀라움, 유연성에서 전문가 평가 인간 응답과 비교한다.
- AUT 응답의 창의성 측정 지표로서 자동화된 의미 거리 측정치를 평가한다.
- AI와 인간의 창의성, 창의성 평가의 의미, 창의성 정의에 대한 시사점을 논의한다.
제안 방법
- Guilford의 Alternative Uses Test 항목에 대해 GPT-3를 적용하여 응답을 생성한다.
- 전문가 평가에서 독창성, 유용성, 놀라움에 대해 인간의 참여 응답과 GPT-3 출력물을 비교한다.
- 항목에 걸친 GPT-3와 인간 간 응답 세트의 유연성을 평가한다.
- 응답과 AUT 객체 간의 의미 거리 기반 자동 창의성 지표를 평가한다.
- AI와 인간의 창의성 해석 및 창의성 테스트의 해석에 대한 논의를 제공한다.
실험 결과
연구 질문
- RQ1GPT-3의 AUT 응답은 전문 평가를 거친 인간 응답과 독창성, 유용성, 놀라움 면에서 어떻게 비교되는가?
- RQ2GPT-3 응답이 AUT 과제에서 인간 응답과 유연성이 비슷한가?
- RQ3의미 거리 기반 자동 창의성 측정치가 인간 평가와 비교하여 AUT 응답에 얼마나 효과적인가?
- RQ4이러한 결과가 AI 창의성과 현재의 창의성 테스트 패러다임에 대해 무엇을 시사하는가?
주요 결과
- 사람은 전문가 평가에 따르면 창의적 산출에서 일반적으로 GPT-3를 능가한다.
- GPT-3는 다소의 창의적 가능성을 보이지만 독창성, 유용성, 놀라움에서 인간 응답에 뒤처진다.
- 향후 이 과제에서 GPT-3가 격차를 줄일 가능성이 있다.
- 본 연구는 인간과 AI의 창의성 및 창의성 테스트 해석에 관한 논의에 정보를 제공한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.