Skip to main content
QUICK REVIEW

[论文解读] Fact-checking information from large language models can decrease headline discernment

Matthew DeVerna, Harry Yaojun Yan|arXiv (Cornell University)|Aug 21, 2023
Misinformation and Its Impacts被引用 4
一句话总结

本研究调查了大型语言模型(LLM)生成的事实核查信息对用户辨别政治新闻准确性及分享意愿的影响。尽管LLM正确驳斥了虚假标题,但它却降低了用户对被错误标记为虚假的正确标题的信任度,并在不确定时增加了对虚假标题的信念,凸显了AI驱动的事实核查系统可能带来的意外危害。

ABSTRACT

Fact checking can be an effective strategy against misinformation, but its implementation at scale is impeded by the overwhelming volume of information online. Recent artificial intelligence (AI) language models have shown impressive ability in fact-checking tasks, but how humans interact with fact-checking information provided by these models is unclear. Here, we investigate the impact of fact-checking information generated by a popular large language model (LLM) on belief in, and sharing intent of, political news headlines in a preregistered randomized control experiment. Although the LLM accurately identifies most false headlines (90%), we find that this information does not significantly improve participants' ability to discern headline accuracy or share accurate news. In contrast, viewing human-generated fact checks enhances discernment in both cases. Subsequent analysis reveals that the AI fact-checker is harmful in specific cases: it decreases beliefs in true headlines that it mislabels as false and increases beliefs in false headlines that it is unsure about. On the positive side, AI fact-checking information increases the sharing intent for correctly labeled true headlines. When participants are given the option to view LLM fact checks and choose to do so, they are significantly more likely to share both true and false news but only more likely to believe false headlines. Our findings highlight an important source of potential harm stemming from AI applications and underscore the critical need for policies to prevent or mitigate such unintended consequences.

研究动机与目标

  • 考察用户在政治新闻背景下对大型语言模型(LLM)生成的事实核查信息的反应。
  • 评估LLM事实核查是否能提升用户辨别新闻标题真伪的能力。
  • 评估LLM事实核查对参与者分享政治新闻意愿的影响。
  • 探究党派一致性在用户接触LLM生成的事实核查后,如何影响其信念与分享行为。
  • 识别在现实信息生态系统中部署LLM进行自动化事实核查的意外后果。

提出的方法

  • 通过预先注册的随机对照实验,招募N=1,548名美国参与者,评估LLM事实核查对新闻感知的因果影响。
  • 参与者接触了40则真实政治新闻,其中一半为真,一半为假,并在党派倾向上保持平衡(支持民主或共和党)。
  • 参与者被分配至“信念”和“分享”组,并被随机分配查看或选择不查看由ChatGPT生成的事实核查信息。
  • 采用统计建模方法,包括稳健标准误的多层次回归,分析事实核查暴露、标题真实性及党派一致性对信念和分享意愿的影响。
  • 检验事实核查选择状态、标题真实性与党派一致性之间的交互效应,以评估不同影响。
  • 通过使用Bonferroni校正的事后比较,评估各条件下信念与分享意愿斜率的差异。
Figure 1: Experimental design, accuracy, and main effects of the LLM fact-checking intervention. (a) Graphical representation of the experimental design and participant flow. Although two different false claims are shown as examples along with their respective ChatGPT fact-checking information, both
Figure 1: Experimental design, accuracy, and main effects of the LLM fact-checking intervention. (a) Graphical representation of the experimental design and participant flow. Although two different false claims are shown as examples along with their respective ChatGPT fact-checking information, both

实验结果

研究问题

  • RQ1接触LLM生成的事实核查是否能提升用户辨别政治新闻标题准确性的能力?
  • RQ2查看LLM事实核查如何影响参与者分享真实与虚假政治新闻的意愿?
  • RQ3LLM对标题的不确定判断或错误标记是否导致对准确新闻信任度下降或对虚假新闻信念度上升?
  • RQ4党派一致性如何调节LLM事实核查对信念与分享意愿的影响?
  • RQ5在公共信息生态系统中部署LLM进行自动化事实核查的意外后果是什么?

主要发现

  • LLM正确驳斥了虚假标题,但其事实核查并未显著提升参与者辨别标题准确性的能力。
  • LLM降低了对被错误标记为虚假的正确标题的信念,表明其对真实信息信任度造成了有害影响。
  • 当LLM对标题真实性不确定时,其事实核查反而增加了对虚假标题的信念,表明存在放大虚假信息的风险。
  • 当参与者选择查看LLM事实核查时,他们更可能分享真实和虚假新闻,但仅对虚假新闻更可能产生信念。
  • 选择不查看事实核查的参与者中,党派不一致对真实标题分享意愿有显著负面影响,表明不查看者对党派倾向更为敏感。
  • 当参与者选择查看事实核查时,党派一致性对分享意愿无显著影响,表明LLM事实核查可能在分享行为中覆盖了党派偏见。
Figure 2: Effects of LLM fact-checking information on headline belief and sharing intent, contingent on headline veracity and fact check judgment. Each panel shows the proportion of participants in the control (circles) and forced (triangles) conditions who (a) believed or (b) were willing to share
Figure 2: Effects of LLM fact-checking information on headline belief and sharing intent, contingent on headline veracity and fact check judgment. Each panel shows the proportion of participants in the control (circles) and forced (triangles) conditions who (a) believed or (b) were willing to share

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。