[论文解读] MediaRank: Computational Ranking of Online News Sources
MediaRank 提出了一套完全自动化且可解释的系统,利用四种计算指标——同行声誉、报道偏见与广度、财务压力以及受欢迎程度——对全球超过50,000家在线新闻源进行排名。该系统与现有专家排名具有较强的显著相关性(斯皮尔曼等级相关系数 ρ = 0.57,p < 0.05,适用于34/35组专家排名),同时引入了新颖的方法来检测社交机器人和激进的广告投放行为。
In the recent political climate, the topic of news quality has drawn attention both from the public and the academic communities. The growing distrust of traditional news media makes it harder to find a common base of accepted truth. In this work, we design and build MediaRank (www.media-rank.com), a fully automated system to rank over 50,000 online news sources around the world. MediaRank collects and analyzes one million news webpages and two million related tweets everyday. We base our algorithmic analysis on four properties journalists have established to be associated with reporting quality: peer reputation, reporting bias / breadth, bottomline financial pressure, and popularity. Our major contributions of this paper include: (i) Open, interpretable quality rankings for over 50,000 of the world's major news sources. Our rankings are validated against 35 published news rankings, including French, German, Russian, and Spanish language sources. MediaRank scores correlate positively with 34 of 35 of these expert rankings. (ii) New computational methods for measuring influence and bottomline pressure. To the best of our knowledge, we are the first to study the large-scale news reporting citation graph in-depth. We also propose new ways to measure the aggressiveness of advertisements and identify social bots, establishing a connection between both types of bad behavior. (iii) Analyzing the effect of media source bias and significance. We prove that news sources cite others despite different political views in accord with quality measures. However, in four English-speaking countries (US, UK, Canada, and Australia), the highest ranking sources all disproportionately favor left-wing parties, even when the majority of news sources exhibited conservative slants.
研究动机与目标
- 开发一种可自动扩展的全球规模在线新闻源排名系统。
- 利用计算信号识别并衡量新闻媒体的关键质量指标。
- 在多种语言和区域的35个专家整理的新闻排名中验证该系统。
- 检测并量化财务压力以及误导性行为(如由机器人驱动的流量和激进广告)的影响。
- 分析新闻源排名中的政治偏见,并评估其对质量感知的影响。
提出的方法
- 从新闻文章构建报道引用图,并计算 PageRank 分数以衡量同行声誉。
- 通过分析大规模新闻语料中对左翼与右翼政治人物的情感差异,量化政治偏见。
- 通过统计独特名人提及次数来衡量报道广度,以评估主题多样性。
- 通过网络行为分析检测社交机器人,并通过投放位置和密度指标量化广告激进程度。
- 将各项信号整合为综合的 MediaRank 得分,归一化至 [0,1] 区间,并通过各信号的百分位数实现可解释性。
- 使用斯皮尔曼等级相关系数和显著性检验,将排名结果与35组专家排名进行对比验证。
实验结果
研究问题
- RQ1计算信号在多种语言和区域中,能否有效预测经专家验证的新闻源质量?
- RQ2同行声誉、偏见、财务压力和受欢迎程度在多大程度上独立预测媒体质量?
- RQ3能否通过自动化手段检测社交机器人和激进广告投放,作为媒体完整性的可靠代理指标?
- RQ4为何在英语国家中,排名靠前的新闻源在整体媒体格局呈现保守倾向的背景下,却明显偏向左翼政治观点?
- RQ5信号组合(如高声誉但广告行为差)如何影响最终排名?
主要发现
- MediaRank 得分与35组专家排名中的34组呈正相关,平均斯皮尔曼等级相关系数为 0.57(24组排名具有 p < 0.05 的统计显著性)。
- 在24组具有统计显著性的专家排名中,该系统的平均排名质量得分为 0.63,表明与人类专家评估高度一致。
- 具有高同行声誉和广泛主题覆盖的新闻源始终排名更高,即使其政治偏见中等。
- 《每日邮报》在声誉和广度方面得分较高,但因广告投放过于激进而被下调,表明系统对财务压力信号具有高度敏感性。
- 在美国、英国、加拿大和澳大利亚,排名靠前的新闻源始终偏向左翼政党,表明尽管整体媒体趋势偏保守,但在高质量新闻源的选择中仍存在系统性偏差。
- 该系统成功识别出社交机器人,并发现其存在与激进广告之间存在显著关联,首次建立了协同性非真实行为与变现压力之间的新型联系。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。