[Paper Review] News Verifiers Showdown: A Comparative Performance Evaluation of ChatGPT 3.5, ChatGPT 4.0, Bing AI, and Bard in News Fact-Checking
The paper evaluates four prominent LLMs (GPT-3.5, GPT-4, Bard, and Bing AI) on 100 fact-checked news items, categorizing responses as True, False, or Partially True/False, and compares them to independent verifications.
This study aimed to evaluate the proficiency of prominent Large Language Models (LLMs), namely OpenAI's ChatGPT 3.5 and 4.0, Google's Bard(LaMDA), and Microsoft's Bing AI in discerning the truthfulness of news items using black box testing. A total of 100 fact-checked news items, all sourced from independent fact-checking agencies, were presented to each of these LLMs under controlled conditions. Their responses were classified into one of three categories: True, False, and Partially True/False. The effectiveness of the LLMs was gauged based on the accuracy of their classifications against the verified facts provided by the independent agencies. The results showed a moderate proficiency across all models, with an average score of 65.25 out of 100. Among the models, OpenAI's GPT-4.0 stood out with a score of 71, suggesting an edge in newer LLMs' abilities to differentiate fact from deception. However, when juxtaposed against the performance of human fact-checkers, the AI models, despite showing promise, lag in comprehending the subtleties and contexts inherent in news information. The findings highlight the potential of AI in the domain of fact-checking while underscoring the continued importance of human cognitive skills and the necessity for persistent advancements in AI capabilities. Finally, the experimental data produced from the simulation of this work is openly available on Kaggle.
Motivation & Objective
- Assess the ability of leading LLMs to distinguish truth from deception in news items using black box testing.
- Compare four major LLMs against independently verified fact checks.
- Quantify overall accuracy and contextual strengths/weaknesses of AI-based fact-checking.
- Provide open data for reproducibility via Kaggle.
Proposed method
- Black box evaluation of four LLMs on 100 fact-checked news items from independent agencies.
- Responses classified into True, False, and Partially True/False categories.
- Accuracy measured as agreement with independent verifications.
- Experimental data made openly available on Kaggle.
Experimental results
Research questions
- RQ1How accurately can each model classify news items as True, False, or Partially True/False?
- RQ2Which model shows the best overall performance in this setup?
- RQ3How do AI model performances compare to human fact-checkers in this dataset?
- RQ4What are the limitations and contexts where AI models struggle in news fact-checking?
Key findings
- Average accuracy across models is 65.25 out of 100.
- GPT-4.0 achieves the highest score at 71.
- All models show moderate proficiency and lag behind human fact-checkers in grasping subtleties and context.
- AI demonstrates potential for fact-checking but requires continued AI capability improvements and human oversight.
- Experimental data from the study is openly available on Kaggle.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.