Skip to main content
QUICK REVIEW

[Paper Review] Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor

Ahmed Sharshar, Hosam Elgendy|arXiv (Cornell University)|Mar 18, 2026
Humor Studies and Applications0 citations
TL;DR

Introduces a multimodal, multilingual benchmark to detect harmful humor across text, image, and video in English and Arabic, including explicit and implicit harm, and evaluates SOTA open/closed models.

ABSTRACT

Dark humor often relies on subtle cultural nuances and implicit cues that require contextual reasoning to interpret, posing safety challenges that current static benchmarks fail to capture. To address this, we introduce a novel multimodal, multilingual benchmark for detecting and understanding harmful and offensive humor. Our manually curated dataset comprises 3,000 texts and 6,000 images in English and Arabic, alongside 1,200 videos that span English, Arabic, and language-independent (universal) contexts. Unlike standard toxicity datasets, we enforce a strict annotation guideline: distinguishing Safe jokes from Harmful ones, with the latter further classified into Explicit (overt) and Implicit (Covert) categories to probe deep reasoning. We systematically evaluate state-of-the-art (SOTA) open and closed-source models across all modalities. Our findings reveal that closed-source models significantly outperform open-source ones, with a notable difference in performance between the English and Arabic languages in both, underscoring the critical need for culturally grounded, reasoning-aware safety alignment. Warning: this paper contains example data that may be offensive, harmful, or biased.

Motivation & Objective

  • Address the gap in safety evaluation for implicitly harmful humor that requires cultural and contextual reasoning.
  • Create a manually curated, cross-modal dataset covering text, images, and videos in English and Arabic (plus universal video context).
  • Evaluate open- and closed-source LLMs/VLMs and video LLMs on a unified harm-detection task.
  • Investigate language-specific weaknesses and the need for culturally grounded safety alignment.

Proposed method

  • Curate 3,000 textual jokes, 6,005 memes/images, and 1,202 short videos across English, Arabic, and universal content with harm labels.
  • Annotate each item as Safe, Harmful with sublabels Explicit or Implicit using majority voting.
  • Evaluate a mix of closed-source (GPT-5.2/4o, Gemini) and open-source (DeepSeek-Reasoner, Qwen, LLaMA-based) models across modalities with binary Harmful vs Safe and recall per Explicit/Implicit.

Experimental results

Research questions

  • RQ1How well can current models detect harmful humor across text, images, and videos in English and Arabic?
  • RQ2Do models show a gap in detecting implicit (contextual) harm versus explicit harm, and is this gap language-dependent?
  • RQ3What is the relative performance of open-source versus closed-source models in multilingual and multimodal harmful-humor detection?
  • RQ4To what extent does language (English vs Arabic) affect safety alignment in multimodal humor understanding?

Key findings

  • Closed-source models generally outperform open-source ones across modalities and languages.
  • There is a notable performance drop from English to Arabic, especially for implicit harm detection.
  • Explicit harm is detected more reliably than implicit harm, with larger gaps in Arabic for many models.
  • Video and image modalities show strong English bias, with Universal content generally underperforming compared to English but over Arabic in some cases.
  • Gemini-2.5-Pro often provides the most balanced performance across modalities, languages, and implicit/explicit harm detection.
  • Open-source models tend to exhibit a safe bias or struggle with multimodal cues, impacting their recall for actual harmful content.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.