[论文解读] A Survey on Predicting the Factuality and the Bias of News Media
本综述提出了一种统一框架,通过利用新闻文章文本以外的多源、多模态数据,联合建模媒体事实性与政治偏见。综述回顾了源级别画像的最先进技术,强调了数据稀缺性和动态评分等挑战,并展望了非英语媒体、视频新闻以及媒体素养实际部署的未来方向。
The present level of proliferation of fake, biased, and propagandistic content online has made it impossible to fact-check every single suspicious claim or article, either manually or automatically. Thus, many researchers are shifting their attention to higher granularity, aiming to profile entire news outlets, which makes it possible to detect likely "fake news" the moment it is published, by simply checking the reliability of its source. Source factuality is also an important element of systems for automatic fact-checking and "fake news" detection, as they need to assess the reliability of the evidence they retrieve online. Political bias detection, which in the Western political landscape is about predicting left-center-right bias, is an equally important topic, which has experienced a similar shift towards profiling entire news outlets. Moreover, there is a clear connection between the two, as highly biased media are less likely to be factual; yet, the two problems have been addressed separately. In this survey, we review the state of the art on media profiling for factuality and bias, arguing for the need to model them jointly. We further discuss interesting recent advances in using different information sources and modalities, which go beyond the text of the articles the target news outlet has published. Finally, we discuss current challenges and outline future research directions.
研究动机与目标
- 回顾并综合近期对新闻媒体整体在事实性与政治偏见方面进行画像的研究进展。
- 论证事实性与偏见的联合建模,鉴于二者存在强烈相互依赖性。
- 识别并分析数据可得性、标注变异性及模型鲁棒性等方面的关键挑战。
- 探索除文章文本外的多样化信息来源,包括社交媒体、维基百科和元数据。
- 概述非英语媒体、视频新闻及面向终端用户的实用工具的未来研究方向。
提出的方法
- 综述综合了基于 Media Bias/Fact Check 和 AllSides 等来源的黄金标准标签的媒体画像现有研究。
- 评估结合文本内容、社交媒体活动、URL 结构和元数据以估计媒体可靠性的方法。
- 分析联合预测事实性与偏见的多任务序数回归模型,使用序数尺度进行建模。
- 研究利用间接信号(如维基百科提及、Twitter 账号和流量数据)推断媒体可信度的方法。
- 评估预训练多语言模型与迁移学习在低资源语言环境中的应用。
- 讨论新闻片段的端到端视频分析技术,包括转录与视觉模态的融合。
实验结果
研究问题
- RQ1如何通过联合建模事实性与偏见来提升对不可靠新闻源的检测能力?
- RQ2除文章内容外,哪些非文本、非文章数据源(如社交媒体、元数据)可增强媒体画像?
- RQ3媒体评分的动态变化如何影响训练数据的可靠性与模型泛化能力?
- RQ4在非英语与非西方政治语境下,扩展媒体画像面临哪些关键挑战?
- RQ5如何实现对视频新闻内容的端到端分析以检测偏见与事实性问题?
主要发现
- 现有媒体事实性与偏见数据集规模较小,通常仅包含数百至数千家新闻媒体。
- 人类标注的事实性与偏见评分本质上是动态的,随时间变化,影响模型评估。
- 通过多任务序数回归实现的事实性与偏见联合建模,相比单任务方法性能更优。
- 当前研究主要局限于英语媒体,针对非英语或多元政党政治体系的研究较少。
- 视频新闻分析仍研究不足,多数研究依赖转录文本而非端到端视频处理。
- 实用工具如 Media Bias/Fact Check 和 AllSides 已存在,但需扩大语言与媒体覆盖范围以实现全球影响。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。