Skip to main content
QUICK REVIEW

[논문 리뷰] ChatGPT is fun, but it is not funny! Humor is still challenging Large Language Models

Sophie Jentzsch, Kristian Kersting|arXiv (Cornell University)|2023. 06. 07.
Topic Modeling인용 수 4
한 줄 요약

이 연구는 프롬프트 기반 실험을 통해 ChatGPT가 유머를 진정으로 이해하고 생성할 수 있는지 조사한다. 농담 생성, 설명, 탐지에 초점을 맞춘 실험에서, 유창하고 맥락 인식 능력이 있는 것처럼 보이지만, 실제로는 원천으로서의 고정된 25개의 기존 농담을 주로 재생산할 뿐이며, 유효하지 않은 농담에 대해도 유사한 설명을 만들어내어, 패턴 매칭을 넘어서는 진정한 유머 이해 능력이 제한되어 있음을 보여준다.

ABSTRACT

Humor is a central aspect of human communication that has not been solved for artificial agents so far. Large language models (LLMs) are increasingly able to capture implicit and contextual information. Especially, OpenAI's ChatGPT recently gained immense public attention. The GPT3-based model almost seems to communicate on a human level and can even tell jokes. Humor is an essential component of human communication. But is ChatGPT really funny? We put ChatGPT's sense of humor to the test. In a series of exploratory experiments around jokes, i.e., generation, explanation, and detection, we seek to understand ChatGPT's capability to grasp and reproduce human humor. Since the model itself is not accessible, we applied prompt-based experiments. Our empirical evidence indicates that jokes are not hard-coded but mostly also not newly generated by the model. Over 90% of 1008 generated jokes were the same 25 Jokes. The system accurately explains valid jokes but also comes up with fictional explanations for invalid jokes. Joke-typical characteristics can mislead ChatGPT in the classification of jokes. ChatGPT has not solved computational humor yet but it can be a big leap toward "funny" machines.

연구 동기 및 목표

  • ChatGPT가 인간의 농담을 진정으로 이해하고 생성하며 설명할 수 있는지 평가하기.
  • 모델이 원천적인 농담을 생성하는지, 아니면 훈련 데이터에서 유래한 사전 존재하는 농담을 단지 재생산하는지 조사하기.
  • 모델이 구조적 및 의미적 특징에 기반해 농담을 탐지할 수 있는지 평가하기.
  • ChatGPT가 농담에 대해 진실된 설명을 제공하는지, 아니면 비유쾌한 농담에 대해서도 이를 위조하는지 검토하기.
  • LLM이 같은 ChatGPT와 같이 표면적 패턴을 넘어서 농담에 대한 깊이 있는 이해를 반영하는 정도를 이해하기.

제안 방법

  • 프롬프트 기반 실험을 수행하여 사전 영향을 피하기 위해 새로운 대화 컨텍스트를 사용하였다.
  • 반복적 프롬프트를 통해 1,008개의 농담을 생성하여 반복성과 다양성을 분석하였다.
  • 유효한 농담과 유효하지 않은 농담을 제공하여 설명의 정확성과 타당성을 평가하기 위해 농담 설명을 평가하였다.
  • 질문-답변 형식, 어법 놀이, 주제 등의 구조적 특성에 기반해 농담 유사 샘플을 분류하여 탐지 능력을 시험하였다.
  • 비유쾌한 농담의 경우를 포함해 일관성, 논리성, 위조 여부를 분석하여 모델의 반응을 평가하였다.
  • 통제된 실험 설계를 통해 모델 행동이 맥락적 영향으로부터 분리되었으며, 본질적 능력에 집중하였다.
Figure 1: Exemplary illustration of a conversation between a human user and an artificial chatbot. The joke is a true response to the presented prompt by ChatGPT.
Figure 1: Exemplary illustration of a conversation between a human user and an artificial chatbot. The joke is a true response to the presented prompt by ChatGPT.

실험 결과

연구 질문

  • RQ1ChatGPT는 얼마나 원천적인 농담을 생성하는가, 아니면 주로 사전 존재하는 고정된 25개의 농담을 반복하는가?
  • RQ2ChatGPT는 농담이 웃기다고 설명할 수 있는가, 아니면 실제로 웃기지 않은 농담에 대해 위조된 설명을 만드는가?
  • RQ3농담처럼 생긴 형식을 가졌지만 실제로는 농담이 아닌 경우, ChatGPT는 농담을 얼마나 잘 탐지하는가?
  • RQ4모델은 표면적 특징(예: 형식, 어법 놀이)에 의존하는가, 아니면 더 깊은 의미적 및 맥락적 농담을 이해하는가?
  • RQ5모델의 행동은 농단에 대한 기본 메커니즘—패턴 매칭인지 진정한 이해인지—를 어떻게 드러내는가?

주요 결과

  • 1,008개의 생성된 농담 중 90퍼센트 이상이 단지 25개의 반복 농담에 속해 있어, 원천적 생성보다는 고정된 농담 풀에 의존하는 것으로 나타났다.
  • 유효한 농담에 대해서는 어법 놀이와 구조적 요소를 식별하여 정확한 설명을 제공하여, 일부 농담 메커니즘을 이해하고 있음을 보여주었다.
  • 비유쾌한 농담에 대해서는 허구적이지만 설득력 있는 설명을 만들어내어, 유사한 추론을 위조하는 경향을 보였다.
  • 다양한 농담 특성(예: 형식, 어법 놀이, 주제)이 함께 존재할수록 농담으로 분류될 가능성이 높아져, 패턴 기반 탐지임을 시사했다.
  • ChatGPT는 단지 외형적인 농담 유사 형식에 속지 않았으며, 형태를 넘어서 내용과 의미를 고려함을 보여주었다.
  • 유창성과 맥락 인식 능력에도 불구하고, 의도적으로 웃기게 만든 원천적 콘텐츠를 생성할 능력이 없어, 계산적 농단 문제를 해결하지 못한 것으로 나타났다.
Figure 2: Modification of top jokes to create joke detection conditions. Below each condition, the percentages of samples are stated that were classified as joke (green), potentially funny (yellow), and not as a joke (red). In condition (A) Minus Wordplay , the comic element, and, therefore, the pun
Figure 2: Modification of top jokes to create joke detection conditions. Below each condition, the percentages of samples are stated that were classified as joke (green), potentially funny (yellow), and not as a joke (red). In condition (A) Minus Wordplay , the comic element, and, therefore, the pun

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.