[Paper Review] Deceptive AI Systems That Give Explanations Are Just as Convincing as Honest AI Systems in Human-Machine Decision Making
This study investigates how deceptive AI explanations affect human judgment in truth discernment tasks. Using a 2×2 randomized experiment with 128 participants, it finds that deceptive AI explanations—falsely claiming true headlines are false or false headlines are true—are just as convincing as honest ones, significantly reducing discernment accuracy regardless of whether the AI is perceived as human or machine. The results highlight a critical vulnerability in human-machine decision-making systems where AI-generated explanations can mislead users even when they are factually incorrect.
The ability to discern between true and false information is essential to making sound decisions. However, with the recent increase in AI-based disinformation campaigns, it has become critical to understand the influence of deceptive systems on human information processing. In experiment (N=128), we investigated how susceptible people are to deceptive AI systems by examining how their ability to discern true news from fake news varies when AI systems are perceived as either human fact-checkers or AI fact-checking systems, and when explanations provided by those fact-checkers are either deceptive or honest. We find that deceitful explanations significantly reduce accuracy, indicating that people are just as likely to believe deceptive AI explanations as honest AI explanations. Although before getting assistance from an AI-system, people have significantly higher weighted discernment accuracy on false headlines than true headlines, we found that with assistance from an AI system, discernment accuracy increased significantly when given honest explanations on both true headlines and false headlines, and decreased significantly when given deceitful explanations on true headlines and false headlines. Further, we did not observe any significant differences in discernment between explanations perceived as coming from a human fact checker compared to an AI-fact checker. Similarly, we found no significant differences in trust. These findings exemplify the dangers of deceptive AI systems and the need for finding novel ways to limit their influence human information processing.
Motivation & Objective
- To examine how deceptive AI explanations affect human ability to discern true from false news in human-machine decision-making contexts.
- To compare the impact of honest versus deceptive explanations on users' truth discernment accuracy, regardless of whether the AI is perceived as human or machine.
- To investigate whether users trust AI fact-checkers as highly as human fact-checkers, and whether this trust is influenced by the honesty of the explanations provided.
- To assess the role of explanation quality and perceived source (human vs. AI) in shaping user judgments and trust in AI-generated fact-checking.
Proposed method
- Conducted a between-subjects randomized 2×2 factorial experiment with 128 participants, assigning them to conditions based on perceived source (human or AI fact-checker) and explanation type (honest or deceptive).
- Generated 14 headlines (7 true, 7 false) with one honest and one deceptive explanation each using GPT-3 (davinci, temp=0.7), prompted to complete 'This is TRUE/FALSE because…'.
- Curated explanations by ranking for semantic similarity and low word repetition, then validated for veracity and logical consistency to ensure balanced, high-quality stimuli.
- Measured participants’ truth discernment on a 7-point Likert scale before and after seeing explanations, calculating weighted discernment accuracy for each condition.
- Collected self-reported trust levels in the fact-checking agent to analyze trust dynamics across conditions.
- Performed t-tests to compare discernment accuracy and trust levels between conditions, testing for significant differences in performance and perception.
Experimental results
Research questions
- RQ1Do deceptive AI explanations reduce users' ability to correctly identify true and false news compared to honest explanations?
- RQ2Does the perceived source of the explanation (human vs. AI fact-checker) affect users' discernment accuracy or trust levels?
- RQ3Are users equally likely to trust AI fact-checkers as they are human fact-checkers, regardless of explanation honesty?
- RQ4Do users show higher discernment accuracy with honest explanations, and does this vary by headline veracity (true vs. false)?
Key findings
- Deceptive explanations significantly reduced discernment accuracy, with mean accuracy dropping to 41.7% on true headlines and 25.8% on false headlines, compared to 83.9% and 79.3% with honest explanations.
- There was no significant difference in discernment accuracy between explanations perceived as coming from a human fact-checker (56.1%) and those from an AI fact-checker (57.6%), p = .427.
- Participants’ trust levels did not differ significantly between human fact-checkers (M = 3.74) and AI fact-checkers (M = 3.72), p = .141, indicating equal perceived reliability.
- Before receiving any explanation, participants were more accurate at identifying false headlines (M = 62.7%) than true headlines (M = 53.1%), p = .000, suggesting a baseline bias toward skepticism of true content.
- With honest explanations, discernment accuracy improved significantly for both true (p = .000) and false headlines (p = .000), confirming the value of truthful explanations.
- The linguistic features (word count, sentiment, grade level, subjectivity) of explanations showed no statistically significant differences between honest and deceptive conditions, indicating that deception was not detectable through surface-level text analysis.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.