[论文解读] In Generative AI we Trust: Can Chatbots Effectively Verify Political Information?
本研究评估了ChatGPT与Bing Chat(现为Microsoft Copilot)在英语、俄语和乌克兰语多语言提示下,对五个敏感话题——新冠疫情、俄罗斯入侵乌克兰、大屠杀、气候变化和LGBTQ+议题——的政治声明进行验证的能力。研究发现,未经微调的ChatGPT在真实性检测中平均准确率达72%,优于Bing Chat的67%,且其表现因语言、话题和信源归属而有显著差异。
This article presents a comparative analysis of the ability of two large language model (LLM)-based chatbots, ChatGPT and Bing Chat, recently rebranded to Microsoft Copilot, to detect veracity of political information. We use AI auditing methodology to investigate how chatbots evaluate true, false, and borderline statements on five topics: COVID-19, Russian aggression against Ukraine, the Holocaust, climate change, and LGBTQ+ related debates. We compare how the chatbots perform in high- and low-resource languages by using prompts in English, Russian, and Ukrainian. Furthermore, we explore the ability of chatbots to evaluate statements according to political communication concepts of disinformation, misinformation, and conspiracy theory, using definition-oriented prompts. We also systematically test how such evaluations are influenced by source bias which we model by attributing specific claims to various political and social actors. The results show high performance of ChatGPT for the baseline veracity evaluation task, with 72 percent of the cases evaluated correctly on average across languages without pre-training. Bing Chat performed worse with a 67 percent accuracy. We observe significant disparities in how chatbots evaluate prompts in high- and low-resource languages and how they adapt their evaluations to political communication concepts with ChatGPT providing more nuanced outputs than Bing Chat. Finally, we find that for some veracity detection-related tasks, the performance of chatbots varied depending on the topic of the statement or the source to which it is attributed. These findings highlight the potential of LLM-based chatbots in tackling different forms of false information in online environments, but also points to the substantial variation in terms of how such potential is realized due to specific factors, such as language of the prompt or the topic.
研究动机与目标
- 评估基于大语言模型(LLM)的聊天机器人在多样化、高风险话题中检测政治声明真实性的能力。
- 探究在高资源语言(英语)与低资源语言(俄语、乌克兰语)之间,性能表现的差异。
- 考察聊天机器人如何理解并应用政治传播概念,如虚假信息、错误信息和阴谋论。
- 分析将声明归因于特定政治或社会主体(如政府、非政府组织、极端组织)对聊天机器人真实性评估的影响。
- 识别在不同提示和语境下,基于LLM的事实核查中系统性偏差与不一致性。
提出的方法
- 采用对比分析方法,使用两种主流的基于大语言模型的聊天机器人:OpenAI的ChatGPT与Microsoft的Bing Chat(重新命名为Copilot)。
- 运用AI审计方法,系统测试聊天机器人在150条陈述上的响应,涵盖五个政治话题:新冠疫情、俄罗斯对乌克兰的军事行动、大屠杀、气候变化和LGBTQ+议题。
- 使用基于定义的提示,评估聊天机器人是否能正确将陈述分类为虚假信息、错误信息或阴谋论。
- 在三种语言中测试性能:英语(高资源语言)、俄语和乌克兰语(低资源语言),以评估语言偏见。
- 通过将相同声明归因于不同政治或社会主体(如政府、非政府组织、极端组织)来模拟信源偏见,以评估响应的一致性。
- 使用真实标签量化所有任务的准确率,并对定性输出进行分析,以捕捉细微差别与一致性。
实验结果
研究问题
- RQ1ChatGPT与Bing Chat在多样化、高敏感度话题中,对政治声明真实性的评估准确率如何?
- RQ2在相同评估任务下,高资源语言(英语)与低资源语言(俄语、乌克兰语)之间的性能表现有何差异?
- RQ3聊天机器人在多大程度上能正确识别并分类政治传播现象,如虚假信息、错误信息和阴谋论?
- RQ4将声明归因于不同政治或社会主体如何影响聊天机器人对其真实性的评估?
- RQ5在基于LLM的事实核查中,性能在话题、语言或信源归属方面是否存在系统性差异?
主要发现
- 未经预训练的ChatGPT在所有语言和话题中,真实性检测的平均准确率达到72%,优于Bing Chat。
- Bing Chat在相同评估任务中的平均准确率为67%,表现较低。
- 在高资源语言与低资源语言之间,性能差异显著,俄语和乌克兰语的准确率明显低于英语。
- 在分类政治传播概念(如虚假信息和阴谋论)时,ChatGPT提供了更具细微差别和语境相关性的回应。
- 性能因话题而异,对气候变迁等事实性话题的准确率较高,而对大屠杀或LGBTQ+议题等情绪化或意识形态敏感话题的准确率较低。
- 信源归属对评估产生可测量影响,聊天机器人对信源的可信度或意识形态存在敏感性,即使声明的事实内容保持不变。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。