Skip to main content
QUICK REVIEW

[論文レビュー] MBIC -- A Media Bias Annotation Dataset Including Annotator Characteristics

Timo Spinde, Lada Rudnitckaia|arXiv (Cornell University)|May 20, 2021
Misinformation and Its Impacts参考文献 11被引用数 11
ひとこと要約

MBICは、政治的傾向やメディア消費習慣を含む、政治的傾向やメディア消費習慣を含む、語と文レベルの偏見言語の両方のアノテーションに加え、アノテーターの特性を詳細に記録した最初のメディアバイアスデータセットを提供する。独自のアノテーションプラットフォームを用いて、1,700件の発言をそれぞれ10名の多様なアノテーターがアノテートした。これにより、異なる背景を持つ人々におけるメディアバイアス認識のより信頼性が高く文脈に配慮した分析が可能になった。

ABSTRACT

Many people consider news articles to be a reliable source of information on current events. However, due to the range of factors influencing news agencies, such coverage may not always be impartial. Media bias, or slanted news coverage, can have a substantial impact on public perception of events, and, accordingly, can potentially alter the beliefs and views of the public. The main data gap in current research on media bias detection is a robust, representative, and diverse dataset containing annotations of biased words and sentences. In particular, existing datasets do not control for the individual background of annotators, which may affect their assessment and, thus, represents critical information for contextualizing their annotations. In this poster, we present a matrix-based methodology to crowdsource such data using a self-developed annotation platform. We also present MBIC (Media Bias Including Characteristics) - the first sample of 1,700 statements representing various media bias instances. The statements were reviewed by ten annotators each and contain labels for media bias identification both on the word and sentence level. MBIC is the first available dataset about media bias reporting detailed information on annotator characteristics and their individual background. The current dataset already significantly extends existing data in this domain providing unique and more reliable insights into the perception of bias. In future, we will further extend it both with respect to the number of articles and annotators per article.

研究の動機と目的

  • 自然言語処理におけるメディアバイアス検出のための、堅牢で代表的かつ多様なデータセットの不足を解消すること。
  • 既存のデータセットの限界を克服するため、バイアス認識に影響を与えるアノテーターの背景要因を組み込むこと。
  • 高品質なメディアバイアスラベルを収集するためのスケーラブルで行列ベースのアノテーション手法を開発すること。
  • より信頼性が高く文脈に配慮したメディアバイアス分析を可能にする、公開可能なデータセットを構築すること。
  • 今後の拡張のための基盤を築くこと。これには、記事数の増加と1記事あたりのアノテーター数の増加が含まれる。

提案手法

  • メディアバイアスアノテーションの行列ベースのクラウドソーシングを支援するため、独自に構築したアノテーションプラットフォームを開発した。
  • 幅広いメディアバイアスの事例をカバーするように、ニュース記事から発言を選定した。
  • 各発言について、語レベルおよび文レベルのバイアスについて、10名の異なるアノテーターがアノテートした。
  • アノテーターは、偏見のある語やフレーズのラベル付けに加え、文全体のバイアス評価も行った。
  • 政治的傾向、メディア消費習慣、人口統計的データを含むアノテーターの特性を、アノテーションと併せて収集した。
  • アノテーターの背景がバイアスラベリング意思決定にどのように影響するかを分析できるように、データセットを構造化した。

実験結果

リサーチクエスチョン

  • RQ1政治的傾向などのアノテーター特性が、ニュースコンテンツにおけるバイアスの特定にどのように影響するか。
  • RQ2異なるメディア文脈において、個々のアノテーターが偏った言語をどの程度異なる認識で捉えているか。
  • RQ3アノテーターの背景データを組み込むことで、メディアバイアス検出モデルの信頼性と解釈可能性が向上するか。
  • RQ4現在の発言およびアノテーターのサンプルは、現実世界のメディアバイアスの多様性をどれほど代表的かつ多様に捉えているか。
  • RQ5複数アノテーターによるラベリングが、バイアスラベリングの一貫性と妥当性にどのような影響を与えるか。

主な発見

  • MBICは、政治的傾向やメディア消費習慣を含む、語と文レベルのバイアスアノテーションに加え、詳細なアノテーター特性を併せ持つ最初のデータセットである。
  • 1,700件の発言がそれぞれ10名の異なるアノテーターによってアノテートされており、高いアノテーター間カバレッジと多様性が保証されている。
  • アノテーターの背景要因がバイアスラベリングに顕著に影響することが判明し、アノテーションの文脈的把握の重要性が浮き彫りになった。
  • 行列ベースのアノテーションアプローチにより、アノテーター間のばらつきを制御しつつ、効率的かつスケーラブルなデータ収集が可能になった。
  • 本データセットは、バイアス認識、モデルの解釈可能性、NLPシステムにおける公平性に関する今後の研究の基盤を提供する。
  • 著者らは、今後の研究において、記事数およびアノテーターの多様性の両面でデータセットを拡張する計画である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。