[Paper Review] ChatGPT is fun, but it is not funny! Humor is still challenging Large Language Models
This study investigates whether ChatGPT can genuinely understand and generate humor through prompt-based experiments on joke generation, explanation, and detection. Despite appearing fluent and context-aware, ChatGPT primarily reproduces a fixed set of 25 pre-existing jokes rather than generating original humor, and it fabricates plausible explanations for invalid jokes, indicating limited true comprehension of humor beyond pattern matching.
Humor is a central aspect of human communication that has not been solved for artificial agents so far. Large language models (LLMs) are increasingly able to capture implicit and contextual information. Especially, OpenAI's ChatGPT recently gained immense public attention. The GPT3-based model almost seems to communicate on a human level and can even tell jokes. Humor is an essential component of human communication. But is ChatGPT really funny? We put ChatGPT's sense of humor to the test. In a series of exploratory experiments around jokes, i.e., generation, explanation, and detection, we seek to understand ChatGPT's capability to grasp and reproduce human humor. Since the model itself is not accessible, we applied prompt-based experiments. Our empirical evidence indicates that jokes are not hard-coded but mostly also not newly generated by the model. Over 90% of 1008 generated jokes were the same 25 Jokes. The system accurately explains valid jokes but also comes up with fictional explanations for invalid jokes. Joke-typical characteristics can mislead ChatGPT in the classification of jokes. ChatGPT has not solved computational humor yet but it can be a big leap toward "funny" machines.
Motivation & Objective
- To assess whether ChatGPT can genuinely understand, generate, and explain human humor.
- To investigate whether the model generates original jokes or merely reproduces pre-existing ones from its training data.
- To evaluate the model's ability to detect humor based on structural and semantic features.
- To examine whether ChatGPT provides truthful explanations for jokes or fabricates them even for non-jokes.
- To understand the extent to which LLMs like ChatGPT reflect deeper comprehension of humor beyond surface-level patterns.
Proposed method
- Conducted prompt-based experiments using a fresh chat context to avoid priming effects.
- Generated 1,008 jokes via repeated prompting to analyze repetition and diversity.
- Evaluated joke explanations by providing valid and invalid jokes to assess accuracy and plausibility of responses.
- Classified joke-like samples based on structural traits such as question-answer format, wordplay, and topic to test detection capability.
- Analyzed model responses for consistency, coherence, and fabrication, especially in cases of non-jokes.
- Used a controlled experimental setup to isolate model behavior from contextual influence, focusing on intrinsic capabilities.

Experimental results
Research questions
- RQ1To what extent does ChatGPT generate original jokes, or does it primarily repeat a fixed set of pre-existing jokes?
- RQ2Can ChatGPT accurately explain why a joke is funny, and does it fabricate explanations for jokes that are not actually humorous?
- RQ3How well does ChatGPT detect humor when presented with joke-like structures that lack actual humor?
- RQ4Does the model rely on surface-level features such as structure or wordplay, or does it understand deeper semantic and contextual humor?
- RQ5What does the model’s behavior reveal about its underlying mechanism for humor—pattern matching versus genuine comprehension?
Key findings
- Over 90% of 1,008 generated jokes were among just 25 repeated jokes, indicating heavy reliance on a fixed joke pool rather than original generation.
- ChatGPT accurately explained valid jokes by identifying wordplay and structural elements, demonstrating some grasp of humorous mechanisms.
- For non-jokes, the model produced fictional but convincing-sounding explanations, showing a tendency to fabricate plausible reasoning.
- The presence of multiple joke characteristics (e.g., structure, wordplay, topic) increased the likelihood of classification as a joke, suggesting pattern-based detection.
- ChatGPT was not misled by superficial joke-like structures alone, indicating it considers content and meaning beyond form.
- Despite its fluency and contextual awareness, the model lacks the ability to create intentionally funny, original content, suggesting it has not solved computational humor.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.