Skip to main content
QUICK REVIEW

[Paper Review] Fakes of Varying Shades: How Warning Affects Human Perception and Engagement Regarding LLM Hallucinations

Mahjabin Nahar, Haeseung Seo|arXiv (Cornell University)|Apr 4, 2024
Risk Perception and Management11 citations
TL;DR

The study shows that humans rank genuine content as most accurate, minor hallucinations next, major hallucinations least, and that warnings reduce perceived accuracy of hallucinations without hurting perceptions of genuine content; warnings also increase dislikes but do not affect likes or shares.

ABSTRACT

The widespread adoption and transformative effects of large language models (LLMs) have sparked concerns regarding their capacity to produce inaccurate and fictitious content, referred to as `hallucinations'. Given the potential risks associated with hallucinations, humans should be able to identify them. This research aims to understand the human perception of LLM hallucinations by systematically varying the degree of hallucination (genuine, minor hallucination, major hallucination) and examining its interaction with warning (i.e., a warning of potential inaccuracies: absent vs. present). Participants (N=419) from Prolific rated the perceived accuracy and engaged with content (e.g., like, dislike, share) in a Q/A format. Participants ranked content as truthful in the order of genuine, minor hallucination, and major hallucination, and user engagement behaviors mirrored this pattern. More importantly, we observed that warning improved the detection of hallucination without significantly affecting the perceived truthfulness of genuine content. We conclude by offering insights for future tools to aid human detection of hallucinations. All survey materials, demographic questions, and post-session questions are available at: https://github.com/MahjabinNahar/fakes-of-varying-shades-survey-materials

Motivation & Objective

  • Understand how untrained evaluators perceive accuracy of LLM-generated content with varying degrees of hallucination (genuine, minor, major).
  • Examine the effect of warning on perceived accuracy and engagement (like, dislike, share) for genuine and hallucinated content.
  • Investigate whether warning alters engagement behaviors and whether effects differ by hallucination level.

Proposed method

  • Generate three response types (genuine, minor hallucination, major hallucination) for 54 questions from TruthfulQA using GPT-3.5-Turbo.
  • Use a 2 (warning vs. control) x 3 (genuine, minor, major) mixed design with a Latin-square presentation of 18 items per group.
  • Measure perceived accuracy on a 5-point scale and collect engagement actions (like, dislike, share) before accuracy ratings.
  • Include a warning tag in the WARN condition: "The responses may contain inaccurate information about people, places, or facts."
  • Recruit 419 Prolific participants (US-based) and perform ANOVAs to test effects and interactions.

Experimental results

Research questions

  • RQ1RQ1: How do untrained evaluators perceive the accuracy of genuine vs. minor vs. major hallucinations, and does warning affect these perceptions?
  • RQ2RQ2: How do untrained evaluators engage (like, dislike, share) with genuine vs. minor vs. major hallucinations, and does warning affect these engagement patterns?

Key findings

  • Content is perceived as accurate in the order: genuine > minor hallucination > major hallucination.
  • Warning lowers perceived accuracy for minor and major hallucinations but does not affect genuine content.
  • Warning increases dislikes for hallucinated content but does not significantly affect likes or shares.
  • Engagement follows accuracy: like and share are higher for genuine content, with minor and major hallucinations receiving progressively fewer engagements.
  • Dislike is higher for hallucinations, especially major ones, and correlates with perceived inaccuracy.
  • Correlations between perceived accuracy and engagement grow stronger with higher hallucination levels.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.