Skip to main content
QUICK REVIEW

[论文解读] Selection Bias in News Coverage: Learning it, Fighting it

Dylan Bourgeois, Jérémie Rappaz|arXiv (Cornell University)|Apr 16, 2019
Media Influence and Politics参考文献 24被引用 10
一句话总结

本文提出一种个性化、基于因子分解的方法,通过从事件报道数据中学习媒体来源的编辑偏好潜在表征,对新闻报道中的选择性偏差进行建模与缓解。通过采用兼顾多样性的重排序目标,该方法将报道不平等性降低(25个来源的基尼系数从0.79降至0.74),同时保留高影响力事件的报道,促进媒体多样性,且不牺牲新闻价值。

ABSTRACT

News entities must select and filter the coverage they broadcast through their respective channels since the set of world events is too large to be treated exhaustively. The subjective nature of this filtering induces biases due to, among other things, resource constraints, editorial guidelines, ideological affinities, or even the fragmented nature of the information at a journalist's disposal. The magnitude and direction of these biases are, however, widely unknown. The absence of ground truth, the sheer size of the event space, or the lack of an exhaustive set of absolute features to measure make it difficult to observe the bias directly, to characterize the leaning's nature and to factor it out to ensure a neutral coverage of the news. In this work, we introduce a methodology to capture the latent structure of media's decision process on a large scale. Our contribution is multi-fold. First, we show media coverage to be predictable using personalization techniques, and evaluate our approach on a large set of events collected from the GDELT database. We then show that a personalized and parametrized approach not only exhibits higher accuracy in coverage prediction, but also provides an interpretable representation of the selection bias. Last, we propose a method able to select a set of sources by leveraging the latent representation. These selected sources provide a more diverse and egalitarian coverage, all while retaining the most actively covered events.

研究动机与目标

  • 将新闻媒体的选择性偏差建模为偏好学习问题,将媒体来源对事件的报道选择视为对事件的潜在偏好。
  • 基于其报道模式,识别可解释的、数据驱动的新闻来源群体,揭示隐藏的结构性关系,如地理关联与企业隶属关系。
  • 开发一种方法,选择代表性来源子集,以确保更富多样性与平等性的新闻报道,同时保留高影响力事件。
  • 量化并缓解新闻关注度分配中的不平等性,以基尼系数为度量指标,同时不牺牲对重大新闻事件的涵盖。
  • 通过减少回音室效应及少数来源主导公众认知的现象,促进媒体多样性。

提出的方法

  • 该方法将新闻报道建模为用户-项目交互矩阵,其中来源为‘用户’,事件为‘项目’,条目表示报道频率。
  • 应用矩阵分解技术,学习来源与事件的低维潜在表征,捕捉媒体报道选择的潜在偏好结构。
  • 采用个性化排序方法,利用多样性参数β对来源选择进行重排序,平衡报道平等性与高影响力事件的保留。
  • 该方法使用重排序目标,最小化基尼系数(不平等性)的同时,最大化所选来源集合中高报道事件的比例。
  • 潜在空间可解释来源间的关系,如地理聚类与内容分发网络,且无需外部元数据。
  • 该方法在基于GDELT的大规模全球新闻事件与来源数据集上进行评估,性能通过预测准确率与公平性指标衡量。

实验结果

研究问题

  • RQ1能否仅基于观测到的报道数据,将新闻报道中的媒体选择性偏差建模为偏好学习问题?
  • RQ2仅从报道模式中,能否揭示新闻来源之间的潜在结构性关系?
  • RQ3基于学习到的表征,重排序策略在多大程度上能改善新闻报道的多样性与平等性?
  • RQ4所提出的方法在报道公平性与高影响力事件保留之间如何权衡?
  • RQ5学习到的表征能否揭示非显而易见的媒体关系,如企业隶属或内容分发网络?

主要发现

  • 与基线模型相比,所提方法在预测媒体报道方面表现出更高的准确性,表明通过个性化因子分解,新闻选择具有可预测性。
  • 学习到的潜在表征揭示了可解释的、连贯的新闻来源群体,包括地理依赖性与同类型媒体隶属关系,即使缺乏辅助信息亦可识别。
  • 当多样性参数β = 0.5时,25个来源的报道不平等基尼系数从0.79降至0.74,表明关注度不平等性降低。
  • 对于100个选定来源,基尼系数从0.78降至0.68,表明在平等性报道方面有显著改善。
  • 尽管覆盖的独立事件数量减少,但重排序后的来源集合相较于随机或基于热度的选择,保留了更大比例的最常被讨论(高影响力)事件。
  • 该方法成功识别出非显而易见的关系,如广播内容分发网络与企业关联网络,增强了媒体生态系统的透明度。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。