Skip to main content
QUICK REVIEW

[Paper Review] In Generative AI we Trust: Can Chatbots Effectively Verify Political Information?

Elizaveta Kuznetsova, Mykola Makhortykh|arXiv (Cornell University)|Dec 20, 2023
Misinformation and Its Impacts4 citations
TL;DR

This study evaluates the ability of ChatGPT and Bing Chat (now Microsoft Copilot) to verify political claims across five sensitive topics—COVID-19, Russian invasion of Ukraine, the Holocaust, climate change, and LGBTQ+ issues—using multilingual prompts in English, Russian, and Ukrainian. It finds that ChatGPT achieves 72% average accuracy in veracity detection without fine-tuning, outperforming Bing Chat’s 67%, with significant variation based on language, topic, and source attribution.

ABSTRACT

This article presents a comparative analysis of the ability of two large language model (LLM)-based chatbots, ChatGPT and Bing Chat, recently rebranded to Microsoft Copilot, to detect veracity of political information. We use AI auditing methodology to investigate how chatbots evaluate true, false, and borderline statements on five topics: COVID-19, Russian aggression against Ukraine, the Holocaust, climate change, and LGBTQ+ related debates. We compare how the chatbots perform in high- and low-resource languages by using prompts in English, Russian, and Ukrainian. Furthermore, we explore the ability of chatbots to evaluate statements according to political communication concepts of disinformation, misinformation, and conspiracy theory, using definition-oriented prompts. We also systematically test how such evaluations are influenced by source bias which we model by attributing specific claims to various political and social actors. The results show high performance of ChatGPT for the baseline veracity evaluation task, with 72 percent of the cases evaluated correctly on average across languages without pre-training. Bing Chat performed worse with a 67 percent accuracy. We observe significant disparities in how chatbots evaluate prompts in high- and low-resource languages and how they adapt their evaluations to political communication concepts with ChatGPT providing more nuanced outputs than Bing Chat. Finally, we find that for some veracity detection-related tasks, the performance of chatbots varied depending on the topic of the statement or the source to which it is attributed. These findings highlight the potential of LLM-based chatbots in tackling different forms of false information in online environments, but also points to the substantial variation in terms of how such potential is realized due to specific factors, such as language of the prompt or the topic.

Motivation & Objective

  • To assess the capacity of large language model (LLM)-based chatbots to detect the truthfulness of political claims across diverse, high-stakes topics.
  • To investigate how performance varies between high- and low-resource languages, including English, Russian, and Ukrainian.
  • To examine how chatbots interpret and apply political communication concepts such as disinformation, misinformation, and conspiracy theories.
  • To analyze the impact of source attribution—assigning claims to specific political or social actors—on the chatbots’ veracity evaluations.
  • To identify systematic biases and inconsistencies in LLM-based fact-checking across different prompts and contexts.

Proposed method

  • Conducted a comparative analysis using two prominent LLM-based chatbots: OpenAI’s ChatGPT and Microsoft’s Bing Chat (rebranded as Copilot).
  • Employed an AI auditing methodology to systematically test chatbot responses on 150 statements across five political topics: COVID-19, Russian aggression in Ukraine, the Holocaust, climate change, and LGBTQ+ debates.
  • Used definition-oriented prompts to evaluate whether chatbots could correctly classify statements as disinformation, misinformation, or conspiracy theories.
  • Tested performance across three languages: English (high-resource), Russian, and Ukrainian (low-resource) to assess linguistic bias.
  • Modeled source bias by attributing identical claims to different political or social actors (e.g., government, NGO, extremist group) to evaluate response consistency.
  • Quantified accuracy across all tasks using ground-truth labels and analyzed qualitative outputs for nuance and consistency.

Experimental results

Research questions

  • RQ1How accurately can ChatGPT and Bing Chat assess the veracity of political claims across diverse, high-sensitivity topics?
  • RQ2How does performance differ between high-resource (English) and low-resource (Russian, Ukrainian) languages in the same evaluation tasks?
  • RQ3To what extent can chatbots correctly identify and classify political communication phenomena such as disinformation, misinformation, and conspiracy theories?
  • RQ4How does attributing a claim to different political or social actors influence the chatbot’s evaluation of its truthfulness?
  • RQ5What are the systematic variations in performance based on topic, language, or source attribution in LLM-based fact-checking?

Key findings

  • ChatGPT achieved an average accuracy of 72% in veracity detection across all languages and topics without pre-training, outperforming Bing Chat.
  • Bing Chat demonstrated lower performance with an average accuracy of 67% across the same evaluation tasks.
  • Significant disparities emerged in performance between high- and low-resource languages, with accuracy dropping notably in Russian and Ukrainian compared to English.
  • ChatGPT provided more nuanced and contextually appropriate responses when classifying political communication concepts such as disinformation and conspiracy theories.
  • Performance varied significantly depending on the topic, with higher accuracy on factual topics like climate change and lower accuracy on emotionally charged or ideologically sensitive topics like the Holocaust or LGBTQ+ debates.
  • Source attribution had a measurable impact on evaluations, with chatbots showing sensitivity to the perceived credibility or ideology of the source, even when the claim’s factual content remained unchanged.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.