[论文解读] MBIC -- A Media Bias Annotation Dataset Including Annotator Characteristics
MBIC 是首个包含词语级和句子级偏见语言标注以及详细标注者特征(如政治倾向和媒体消费习惯)的媒体偏见数据集。通过自定义标注平台,作者收集了由十位不同标注者分别标注的1,700条语句,实现了对不同背景人群媒体偏见感知的更可靠、更具上下文意识的分析。
Many people consider news articles to be a reliable source of information on current events. However, due to the range of factors influencing news agencies, such coverage may not always be impartial. Media bias, or slanted news coverage, can have a substantial impact on public perception of events, and, accordingly, can potentially alter the beliefs and views of the public. The main data gap in current research on media bias detection is a robust, representative, and diverse dataset containing annotations of biased words and sentences. In particular, existing datasets do not control for the individual background of annotators, which may affect their assessment and, thus, represents critical information for contextualizing their annotations. In this poster, we present a matrix-based methodology to crowdsource such data using a self-developed annotation platform. We also present MBIC (Media Bias Including Characteristics) - the first sample of 1,700 statements representing various media bias instances. The statements were reviewed by ten annotators each and contain labels for media bias identification both on the word and sentence level. MBIC is the first available dataset about media bias reporting detailed information on annotator characteristics and their individual background. The current dataset already significantly extends existing data in this domain providing unique and more reliable insights into the perception of bias. In future, we will further extend it both with respect to the number of articles and annotators per article.
研究动机与目标
- 为自然语言处理中的媒体偏见检测解决缺乏稳健、具有代表性且多样化的数据集的问题。
- 通过引入影响偏见感知的标注者背景因素,克服现有数据集的局限性。
- 开发一种可扩展的基于矩阵的标注方法,以收集高质量的媒体偏见标签。
- 创建一个公开可用的数据集,以实现对媒体偏见更可靠、更具上下文意识的分析。
- 为未来扩展奠定基础,包括增加文章数量和每篇文章的标注者数量。
提出的方法
- 开发了自定义标注平台,以支持基于矩阵的众包媒体偏见标注。
- 从新闻文章中选取语句,以代表广泛的媒体偏见实例。
- 每条语句由十位不同的标注者在词语和句子两个层级上进行偏见标注。
- 标注者提供了对偏见词语和短语的标签,以及整体句子层级的偏见评估。
- 在标注的同时,收集了标注者的特征信息,包括政治倾向、媒体消费习惯和人口统计学数据。
- 数据集的结构设计使得能够分析个体标注者背景如何影响偏见标注决策。
实验结果
研究问题
- RQ1标注者特征(如政治倾向)如何影响对新闻内容中媒体偏见的识别?
- RQ2不同标注者在不同媒体语境下对偏见语言的感知差异有多大?
- RQ3整合标注者背景数据能否提高媒体偏见检测模型的可靠性和可解释性?
- RQ4当前语句和标注者样本在捕捉现实世界媒体偏见变化方面的代表性与多样性如何?
- RQ5多标注者标注对偏见标注的一致性和有效性有何影响?
主要发现
- MBIC 是首个同时包含媒体偏见标注和详细标注者特征(如政治倾向和媒体消费习惯)的数据集。
- 该数据集包含由十位不同标注者分别标注的1,700条语句,确保了高标注者间覆盖率和多样性。
- 研究发现标注者背景因素显著影响偏见标注,凸显了在标注中考虑上下文的重要性。
- 基于矩阵的标注方法实现了高效且可扩展的数据收集,同时控制了标注者之间的变异性。
- 该数据集为未来关于偏见感知、模型可解释性以及自然语言处理系统公平性的研究奠定了基础。
- 作者计划在未来工作中进一步扩展数据集,包括增加文章数量和提升标注者多样性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。