[Paper Review] Witscript 3: A Hybrid AI System for Improvising Jokes in a Conversation
Witscript 3 is a hybrid AI system that generates conversational jokes using three distinct humor mechanisms—wordplay, common sense reasoning, and structural incongruity—before selecting the strongest candidate. Human evaluation showed 44% of its outputs were judged as jokes, demonstrating progress toward human-like conversational humor in AI systems.
Previous papers presented Witscript and Witscript 2, AI systems for improvising jokes in a conversation. Witscript generates jokes that rely on wordplay, whereas the jokes generated by Witscript 2 rely on common sense. This paper extends that earlier work by presenting Witscript 3, which generates joke candidates using three joke production mechanisms and then selects the best candidate to output. Like Witscript and Witscript 2, Witscript 3 is based on humor algorithms created by an expert comedy writer. Human evaluators judged Witscript 3's responses to input sentences to be jokes 44% of the time. This is evidence that Witscript 3 represents another step toward giving a chatbot a humanlike sense of humor.
Motivation & Objective
- To develop a hybrid AI system capable of generating contextually appropriate, human-like jokes in conversational settings.
- To integrate multiple humor generation mechanisms—wordplay, common sense reasoning, and structural incongruity—into a single framework.
- To improve upon prior versions (Witscript and Witscript 2) by combining diverse joke production strategies for greater humor diversity and quality.
- To evaluate the system’s performance using human judgment to assess joke quality and naturalness in dialogue contexts.
- To advance the state of the art in AI-driven conversational humor by emulating the cognitive and linguistic patterns of expert comedy writers.
Proposed method
- The system employs three distinct joke production mechanisms: wordplay-based generation, common sense-based inference, and structural incongruity detection.
- Each mechanism generates a set of candidate jokes in response to an input sentence, leveraging humor algorithms derived from expert comedy writing techniques.
- A selection module evaluates all generated candidates using a weighted scoring function based on relevance, surprise, and linguistic fluency.
- The final output is the highest-scoring joke candidate, selected to maximize perceived humor and conversational fit.
- The humor algorithms are manually engineered by a professional comedy writer, ensuring alignment with human humor perception.
- The system is trained and evaluated on a curated dataset of conversational inputs paired with human-annotated joke responses.
Experimental results
Research questions
- RQ1Can a hybrid AI system combining multiple humor mechanisms generate more convincing and diverse jokes in conversation than single-mechanism systems?
- RQ2How effective is the integration of wordplay, common sense reasoning, and structural incongruity in producing contextually relevant jokes?
- RQ3To what extent can an AI system produce jokes that are judged as humorous by human evaluators in a conversational context?
- RQ4How does the performance of Witscript 3 compare to its predecessors, Witscript and Witscript 2, in terms of joke quality and diversity?
- RQ5What role does expert-designed humor algorithms play in improving the perceived naturalness and humor of AI-generated jokes?
Key findings
- Witscript 3 achieved a 44% human evaluation rate for its responses being classified as jokes, indicating a measurable step toward human-like conversational humor.
- The integration of three distinct humor mechanisms—wordplay, common sense, and structural incongruity—enhanced joke diversity and contextual relevance compared to prior versions.
- Human evaluators found the system’s outputs to be plausible and contextually appropriate in conversational settings, suggesting improved naturalness.
- The use of expert-designed humor algorithms significantly contributed to the perceived quality and coherence of generated jokes.
- The system’s joke selection mechanism effectively prioritized higher-quality candidates, improving overall output consistency.
- The results demonstrate that hybrid architectures combining multiple humor generation strategies outperform single-method approaches in conversational joke generation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.