Skip to main content
QUICK REVIEW

[Paper Review] Multi-Agent Comedy Club: Investigating Community Discussion Effects on LLM Humor Generation

Shiwei Hong, Lingyao Li|arXiv (Cornell University)|Feb 16, 2026
Humor Studies and Applications0 citations
TL;DR

The paper shows that broadcast community discussion, stored as social memory and retrieved across rounds, improves long-form stand-up humor generation by LLM agents, compared to a no-discussion baseline.

ABSTRACT

Prior work has explored multi-turn interaction and feedback for LLM writing, but evaluations still largely center on prompts and localized feedback, leaving persistent public reception in online communities underexamined. We test whether broadcast community discussion improves stand-up comedy writing in a controlled multi-agent sandbox: in the discussion condition, critic and audience threads are recorded, filtered, stored as social memory, and later retrieved to condition subsequent generations, whereas the baseline omits discussion. Across 50 rounds (250 paired monologues) judged by five expert annotators using A/B preference and a 15-item rubric, discussion wins 75.6% of instances and improves Craft/Clarity (Δ = 0.440) and Social Response (Δ = 0.422), with occasional increases in aggressive humor.

Motivation & Objective

  • Motivate and quantify how public reception signals influence iterative, long-form humor generation.
  • Isolated evaluation of cross-round reception as a conditioning signal separate from within-round revision.
  • Build a controlled sandbox to compare discussion-enabled versus baseline humor generation across rounds.
  • Provide reusable dataset and evaluation protocol for reception-grounded creative generation.

Proposed method

  • Design a closed sandbox with 35 GPT-4o-mini agents (5 performers, 3 critics, 26 audience, 1 host).
  • Manipulate whether post-performance discussion is enabled (g=1) or skipped (g=0).
  • Use a bounded social memory interface that retrieves memory items into performer contexts across rounds.
  • Log and reconstruct discussion threads as memory blocks retrieved via embedding-based similarity scores.
  • Evaluate paired outputs with human raters using forced A/B preference and a 15-item rubric spanning outcomes, craft, and social reception.
  • Employ a fixed topic sequence of 50 rounds; performers write exactly one monologue per round; no within-round revisions.

Experimental results

Research questions

  • RQ1Does broadcast community discussion improve long-form humor generation compared to a baseline without discussion?
  • RQ2What are the craft, clarity, and social reception effects of incorporating reception-grounded conditioning across rounds?
  • RQ3What tradeoffs in humor style or safety accompany discussion-driven improvements?
  • RQ4How stable are the observed effects across rounds and performer personas?

Key findings

  • Discussion-enabled outputs win 75.6% of paired instances (A/B preference).
  • Craft/Clarity gains with discussion: Delta = 0.440 over baseline.
  • Social Response gains with discussion: Delta = 0.422 over baseline.
  • Immediate amusement (Q1) improves with discussion (0.52 delta on average).
  • Memorability (Q12) and Task Attraction (Q15) show positive shifts under discussion.
  • There is a potential shift toward edgier/harmful humor (HarmShift analysis) in some instances.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.