Skip to main content
QUICK REVIEW

[Paper Review] Do You Trust ChatGPT? -- Perceived Credibility of Human and AI-Generated Content

Martin Huschens, Martin Briesch|arXiv (Cornell University)|Sep 5, 2023
Artificial Intelligence in Healthcare and Education18 citations
TL;DR

The study compares perceived credibility of human- vs AI-generated content across three UI conditions and finds similar credibility for both origins, with AI content rated clearer and more engaging.

ABSTRACT

This paper examines how individuals perceive the credibility of content originating from human authors versus content generated by large language models, like the GPT language model family that powers ChatGPT, in different user interface versions. Surprisingly, our results demonstrate that regardless of the user interface presentation, participants tend to attribute similar levels of credibility. While participants also do not report any different perceptions of competence and trustworthiness between human and AI-generated content, they rate AI-generated content as being clearer and more engaging. The findings from this study serve as a call for a more discerning approach to evaluating information sources, encouraging users to exercise caution and critical thinking when engaging with content generated by AI systems.

Motivation & Objective

  • Assess how UI design (ChatGPT, raw text, Wikipedia-style) affects credibility judgments of text excerpts.
  • Compare perceived credibility of human-generated versus LLM-generated content.
  • Examine whether perceptions of competence, trustworthiness, clarity, and engagement differ by origin and UI.
  • Control for participant demographics and related covariates to ensure valid comparisons.

Proposed method

  • Online survey with 606 English-speaking participants recruited via Prolific.
  • Each participant evaluated four topics (Academy Awards, Canada, Malware, US Senate) presented in two origins (human-generated, LLM-generated) and three UI conditions.
  • Credibility measured with 11 items on a 5-point Likert scale across four factors (competence, trustworthiness, clarity, engagement).
  • Confirmatory factor analysis (CFA) validated a four-factor measurement model with reliability and validity checks.
  • Non-parametric tests (Kruskal-Wallis, Wilcoxon) and descriptive visuals used to compare UI conditions and content origin.

Experimental results

Research questions

  • RQ1Does the UI condition (ChatGPT, Raw Text, Wikipedia UI) influence perceived credibility of the text excerpts?
  • RQ2Are there differences in perceived credibility between human-generated and LLM-generated content across UI conditions?
  • RQ3Which dimensions of credibility (competence, trustworthiness, clarity, engagement) are affected by content origin or UI condition?

Key findings

  • UI condition had no substantial effect on credibility perceptions or reading performance across dimensions.
  • Participants showed no differences in competence or trustworthiness between human- and LLM-generated content.
  • LLM-generated text was perceived as clearer and more engaging than human-generated text, with shorter reading times.
  • Within UI groups, clarity and engagement favored LLM-generated content with small-to-moderate effect sizes (clarity: p<0.01, r=0.261; engagement: p<0.01, r=0.134).
  • Reading time was significantly shorter for LLM-generated content (p<0.01, r=0.114).
  • The study highlights a need for critical evaluation and potential labeling of AI-generated information to mitigate misinformation risks.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.