Skip to main content
QUICK REVIEW

[Paper Review] Can ChatGPT-like Generative Models Guarantee Factual Accuracy? On the Mistakes of New Generation Search Engines

Ruochen Zhao, Xingxuan Li|arXiv (Cornell University)|Mar 3, 2023
Artificial Intelligence in Healthcare and Education15 citations
TL;DR

The paper analyzes factual mistakes of AI-powered search engines (Bing and Bard) and argues that ChatGPT-like models cannot guarantee factual accuracy given current limitations, advocating for transparency and grounding improvements.

ABSTRACT

Although large conversational AI models such as OpenAI's ChatGPT have demonstrated great potential, we question whether such models can guarantee factual accuracy. Recently, technology companies such as Microsoft and Google have announced new services which aim to combine search engines with conversational AI. However, we have found numerous mistakes in the public demonstrations that suggest we should not easily trust the factual claims of the AI models. Rather than criticizing specific models or companies, we hope to call on researchers and developers to improve AI models' transparency and factual correctness.

Motivation & Objective

  • Highlight factual grounding failures in AI-powered search demonstrations by Bing and Bard.
  • Illustrate types of factual errors (conflicts with sources, non-existent details, and unreferenced claims).
  • Discuss short- and long-term strategies to improve transparency, source provenance, and factual correctness in conversational models.

Proposed method

  • Systematic review of publicly demonstrated examples from Microsoft Bing and Google Bard demos.
  • Categorization of factual mistakes into three main types: conflicting with sources, not present in sources, and source-inconsistent/unreferenced claims.
  • Comparison of Bing and Bard demonstrations and assessment of transparency and grounding.
  • Discussion of potential remedies including model transparency, confidence reporting, and source-based verification.

Experimental results

Research questions

  • RQ1What kinds of factual mistakes are exhibited by Bing and Bard in demonstrations?
  • RQ2To what extent do these mistakes reflect fundamental grounding issues in ChatGPT-like models?
  • RQ3How do transparency and source citation affect trust in AI-assisted search results?
  • RQ4What short- and long-term approaches could improve factual accuracy in conversational search engines?

Key findings

  • The new Bing demonstration produced fabricated financial data and erroneous comparative tables not supported by original reports.
  • Bing also provided incorrect personal details and time-sensitive information (e.g., nightclub hours) inconsistent with sources.
  • Bard demonstrations contained errors such as incorrect telescope discovery attribution and constellation visibility timing, leading to public stock impact.
  • Both systems showed limitations in factual grounding, with some outputs lacking citations or relying on unreliable sources.
  • The authors observe that Bing’s references were more transparent than Bard’s, enabling easier user fact-checking.
  • The paper argues that current ChatGPT-like models cannot guarantee factual accuracy and emphasizes need for transparency and verifiable grounding.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.