Skip to main content
QUICK REVIEW

[Paper Review] GOAT-Bench: Safety Insights to Large Multimodal Models through Meme-Based Social Abuse

Hongzhan Lin, Ziyang Luo|arXiv (Cornell University)|Jan 3, 2024
Hate Speech and Cyberbullying Detection4 citations
TL;DR

This paper introduces GOAT-Bench, a benchmark of 6,626 memes targeting social abuse detection in multimodal AI models. It evaluates large multimodal models (LMMs) on hatefulness, misogyny, offensiveness, sarcasm, and harmfulness, revealing significant safety gaps despite state-of-the-art performance, especially in detecting implicit abuse.

ABSTRACT

The exponential growth of social media has profoundly transformed how information is created, disseminated, and absorbed, exceeding any precedent in the digital age. Regrettably, this explosion has also spawned a significant increase in the online abuse of memes. Evaluating the negative impact of memes is notably challenging, owing to their often subtle and implicit meanings, which are not directly conveyed through the overt text and image. In light of this, large multimodal models (LMMs) have emerged as a focal point of interest due to their remarkable capabilities in handling diverse multimodal tasks. In response to this development, our paper aims to thoroughly examine the capacity of various LMMs (e.g., GPT-4o) to discern and respond to the nuanced aspects of social abuse manifested in memes. We introduce the comprehensive meme benchmark, GOAT-Bench, comprising over 6K varied memes encapsulating themes such as implicit hate speech, sexism, and cyberbullying, etc. Utilizing GOAT-Bench, we delve into the ability of LMMs to accurately assess hatefulness, misogyny, offensiveness, sarcasm, and harmful content. Our extensive experiments across a range of LMMs reveal that current models still exhibit a deficiency in safety awareness, showing insensitivity to various forms of implicit abuse. We posit that this shortfall represents a critical impediment to the realization of safe artificial intelligence. The GOAT-Bench and accompanying resources are publicly accessible at https://goatlmm.github.io/, contributing to ongoing research in this vital field.

Motivation & Objective

  • To address the growing challenge of detecting subtle, implicit social abuse in memes—especially those involving hate speech, sexism, and cyberbullying—due to their multimodal, context-dependent nature.
  • To systematically evaluate the safety awareness of large multimodal models (LMMs) in identifying nuanced forms of online abuse embedded in memes.
  • To develop a comprehensive, publicly available benchmark (GOAT-Bench) that captures diverse, real-world meme-based abuse across multiple dimensions of harm.
  • To investigate the effectiveness of chain-of-thought (CoT) prompting in improving LMMs’ reasoning and detection accuracy for complex, implicit abuse in memes.
  • To contribute to safer AI development by exposing limitations in current LMMs’ social and ethical reasoning, especially in high-stakes, real-world social media contexts.

Proposed method

  • Constructed GOAT-Bench, a curated dataset of 6,626 memes collected from public sources, including the Facebook Hateful Meme dataset, with explicit annotations for five abuse categories: hatefulness, misogyny, offensiveness, sarcasm, and harmfulness.
  • Employed a multi-stage annotation process with human evaluators trained to identify subtle, implicit forms of abuse, ensuring high inter-annotator agreement and psychological safety through consent, workload caps, and well-being checks.
  • Evaluated a range of state-of-the-art LMMs—including GPT-4V, LLaVA-1.5, InstructBLIP, CogVLM, Qwen-VL, and MiniGPT-4—on the benchmark using both zero-shot and chain-of-thought (CoT) prompting strategies.
  • Designed task-specific evaluation protocols to assess model performance across five distinct abuse detection tasks, enabling granular analysis of model strengths and weaknesses.
  • Implemented rigorous data governance: only text annotations from Facebook memes were included in GOAT-Bench, with original images to be downloaded separately from the Facebook Hateful Meme challenge to comply with licensing.
  • Used standardized metrics to compare model performance across tasks, including accuracy, F1-score, and AUC, with GPT-4V achieving the highest overall performance.

Experimental results

Research questions

  • RQ1To what extent can current large multimodal models (LMMs) accurately detect implicit hate speech and social abuse in memes, particularly when the meaning relies on visual and textual cues combined?
  • RQ2How does the use of chain-of-thought (CoT) prompting affect LMMs’ ability to reason through and correctly classify nuanced, multimodal forms of abuse such as sarcasm and subtle misogyny?
  • RQ3What are the key failure modes of LMMs when confronted with memes that embed harmful content through cultural, political, or ironic references rather than explicit language?
  • RQ4How do different LMM architectures and training data distributions influence performance on the detection of subtle, context-dependent forms of online abuse in memes?
  • RQ5Can a standardized, comprehensive benchmark like GOAT-Bench effectively reveal safety shortcomings in LMMs that are not captured by traditional multimodal benchmarks?

Key findings

  • GPT-4V achieved the highest overall performance across all five abuse detection tasks on GOAT-Bench, outperforming other leading LMMs such as LLaVA-1.5, InstructBLIP, and Qwen-VL.
  • Despite strong performance, even GPT-4V showed significant limitations in detecting implicit abuse, particularly in sarcasm and subtle misogynistic content, indicating a fundamental gap in safety reasoning.
  • Chain-of-thought (CoT) prompting improved detection accuracy for several LMMs, especially in complex reasoning tasks like identifying sarcasm and indirect hate speech, though gains were inconsistent across models.
  • Many LMMs failed to recognize harmful memes that relied on cultural or political references, demonstrating poor generalization to contextually embedded abuse beyond surface-level text or image features.
  • The study revealed that current LMMs often misclassify memes with implicit or ironic abuse as non-offensive, highlighting a critical safety vulnerability in real-world deployment scenarios.
  • The benchmark exposed that existing models lack robust alignment with human social values in detecting nuanced online abuse, especially when visual and textual cues contradict or obscure the intended message.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.