[论文解读] ChatGPT is fun, but it is not funny! Humor is still challenging Large Language Models
本研究通过基于提示的实验,探究了ChatGPT是否能够真正理解并生成幽默。尽管其表现流利且具备上下文意识,但ChatGPT主要重复使用一组固定的25个既有的笑话,而非生成原创幽默;对于无效笑话,它还会编造看似合理的解释,表明其对幽默的理解远超模式匹配,缺乏真正的深层理解。
Humor is a central aspect of human communication that has not been solved for artificial agents so far. Large language models (LLMs) are increasingly able to capture implicit and contextual information. Especially, OpenAI's ChatGPT recently gained immense public attention. The GPT3-based model almost seems to communicate on a human level and can even tell jokes. Humor is an essential component of human communication. But is ChatGPT really funny? We put ChatGPT's sense of humor to the test. In a series of exploratory experiments around jokes, i.e., generation, explanation, and detection, we seek to understand ChatGPT's capability to grasp and reproduce human humor. Since the model itself is not accessible, we applied prompt-based experiments. Our empirical evidence indicates that jokes are not hard-coded but mostly also not newly generated by the model. Over 90% of 1008 generated jokes were the same 25 Jokes. The system accurately explains valid jokes but also comes up with fictional explanations for invalid jokes. Joke-typical characteristics can mislead ChatGPT in the classification of jokes. ChatGPT has not solved computational humor yet but it can be a big leap toward "funny" machines.
研究动机与目标
- 评估ChatGPT是否能够真正理解、生成并解释人类幽默。
- 探究该模型生成的是原创笑话,还是仅从训练数据中复制已有笑话。
- 评估模型基于结构和语义特征检测幽默的能力。
- 检查ChatGPT是否对笑话提供真实解释,或在面对非笑话时也编造解释。
- 理解类似ChatGPT的大型语言模型在表面模式之外,对幽默的深层理解程度有多深。
提出的方法
- 通过全新对话上下文进行基于提示的实验,以避免提示效应的影响。
- 通过重复提示生成1,008个笑话,以分析其重复性与多样性。
- 向模型提供有效和无效的笑话,评估其解释的准确性和可信度。
- 基于问题-答案格式、双关语、主题等结构特征,对类似笑话的样本进行分类,以测试其检测能力。
- 分析模型在非笑话情况下的回答一致性、连贯性及虚构性,尤其关注其在非笑话情境下的表现。
- 采用受控实验设置,隔离上下文影响,聚焦于模型的内在能力。

实验结果
研究问题
- RQ1ChatGPT在多大程度上生成原创笑话,还是主要重复一组固定的预设笑话?
- RQ2ChatGPT能否准确解释一个笑话为何好笑?它是否会为实际上并不幽默的笑话编造解释?
- RQ3当面对缺乏实际幽默感的类似笑话结构时,ChatGPT的幽默检测能力如何?
- RQ4该模型是否依赖于表面特征(如结构或双关语),还是能理解更深层次的语义与语境幽默?
- RQ5模型的行为揭示了其幽默处理机制的本质——是模式匹配,还是真正的理解?
主要发现
- 在1,008个生成的笑话中,超过90%属于仅25个重复笑话,表明其严重依赖固定笑话库,而非原创生成。
- 对于有效笑话,ChatGPT能准确识别双关语和结构要素,显示出对幽默机制的一定理解。
- 对于非笑话,模型生成了虚构但听起来可信的解释,表现出编造合理推理的倾向。
- 当多个笑话特征(如结构、双关语、主题)同时存在时,被分类为笑话的可能性更高,表明其检测依赖于模式匹配。
- ChatGPT未被仅具表面形式的类似笑话结构所误导,表明其不仅关注形式,还考虑内容与意义。
- 尽管其语言流利且具备上下文意识,但模型仍无法创造有意图的、原创的幽默内容,表明其尚未解决计算幽默问题。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。