Skip to main content
QUICK REVIEW

[论文解读] Can ChatGPT-like Generative Models Guarantee Factual Accuracy? On the Mistakes of New Generation Search Engines

Ruochen Zhao, Xingxuan Li|arXiv (Cornell University)|Mar 3, 2023
Artificial Intelligence in Healthcare and Education被引用 15
一句话总结

该论文分析基于AI驱动的搜索引擎(Bing 与 Bard)的事实错误,并指出像 ChatGPT 一样的模型在当前限制下不能保证事实准确性,主张提高透明度与 grounding 的改进。

ABSTRACT

Although large conversational AI models such as OpenAI's ChatGPT have demonstrated great potential, we question whether such models can guarantee factual accuracy. Recently, technology companies such as Microsoft and Google have announced new services which aim to combine search engines with conversational AI. However, we have found numerous mistakes in the public demonstrations that suggest we should not easily trust the factual claims of the AI models. Rather than criticizing specific models or companies, we hope to call on researchers and developers to improve AI models' transparency and factual correctness.

研究动机与目标

  • 突出 Bing 与 Bard 的 AI 驱动搜索演示中的事实 grounding 失败。
  • 说明事实错误的类型(与来源冲突、來源中不存在的细节、未引用的主张)。
  • 讨论在对话模型中提高透明度、来源证明和事实正确性的短期与长期策略。

提出的方法

  • 对 Microsoft Bing 和 Google Bard 演示的公开示例进行系统性回顾。
  • 将事实错误分为三大类:与来源冲突、来源中不存在、以及来源不一致/未引用的主张。
  • 比较 Bing 与 Bard 的演示并评估透明度与 grounding。
  • 讨论包括模型透明度、置信度报告和基于来源的验证在内的潜在 remedy。

实验结果

研究问题

  • RQ1Bing 与 Bard 在演示中表现出哪些类型的事实错误?
  • RQ2这些错误在多大程度上反映了 ChatGPT 类模型的根本 grounding 问题?
  • RQ3透明度和来源引用如何影响对 AI 辅助搜索结果的信任?
  • RQ4哪些短期和长期的方法可以改进对话式搜索引擎的事实准确性?

主要发现

  • 新的 Bing 演示中产生了被原创报道所不支持的捏造财务数据和错误的对比表。
  • Bing 还提供了不正确的个人信息和时效性信息(例如夜总会营业时间),与来源不一致。
  • Bard 演示包含诸如错误的望远镜发现归因和星座可见性时间等错误,造成股市公开影响。
  • 两种系统在事实 grounding 上存在局限性,部分输出缺乏引用或依赖于不可靠来源。
  • 作者观察到 Bing 的引用比 Bard 更透明,便于用户进行事实核查。
  • 论文认为当前的 ChatGPT 类模型不能保证事实准确性,并强调需要透明度和可核实的 grounding。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。