[论文解读] NeMig -- A Bilingual News Collection and Knowledge Graph about Migration
NeMig 引入了一个双语(德语/英语)新闻语料库及知识图谱,聚焦于移民议题,对情感极化、媒体政治倾向、子主题以及经 Wikidata 消歧的实体进行了标注。该数据集包含经过匿名化的用户数据,涵盖社会人口统计与政治属性,支持对推荐系统偏差、信息过滤气泡以及跨语言新闻筛选效应的分析,超越传统准确性评估。
News recommendation plays a critical role in shaping the public's worldviews through the way in which it filters and disseminates information about different topics. Given the crucial impact that media plays in opinion formation, especially for sensitive topics, understanding the effects of personalized recommendation beyond accuracy has become essential in today's digital society. In this work, we present NeMig, a bilingual news collection on the topic of migration, and corresponding rich user data. In comparison to existing news recommendation datasets, which comprise a large variety of monolingual news, NeMig covers articles on a single controversial topic, published in both Germany and the US. We annotate the sentiment polarization of the articles and the political leanings of the media outlets, in addition to extracting subtopics and named entities disambiguated through Wikidata. These features can be used to analyze the effects of algorithmic news curation beyond accuracy-based performance, such as recommender biases and the creation of filter bubbles. We construct domain-specific knowledge graphs from the news text and metadata, thus encoding knowledge-level connections between articles. Importantly, while existing datasets include only click behavior, we collect user socio-demographic and political information in addition to explicit click feedback. We demonstrate the utility of NeMig through experiments on the tasks of news recommenders benchmarking, analysis of biases in recommenders, and news trends analysis. NeMig aims to provide a useful resource for the news recommendation community and to foster interdisciplinary research into the multidimensional effects of algorithmic news curation.
研究动机与目标
- 解决在争议性话题(如移民)上缺乏多语言、主题特定的新闻数据集,且缺乏政治与情感标注的问题。
- 支持对推荐算法偏差与过滤气泡形成机制的研究,超越基于准确性的评估。
- 为媒体筛选效应在政治极化与公众舆论方面的影响提供跨学科研究资源。
- 通过结构化、经 Wikidata 增强的知识图谱,支持知识感知推荐系统的发展。
- 通过提供丰富的用户档案与明确的点击反馈,促进合成用户数据的生成。
提出的方法
- 从涵盖政治光谱的多样化媒体来源中,精心筛选了 7,000 篇德语和 10,000 篇美国新闻文章。
- 通过专家标注与现有分类方案,对文章进行情感极化与媒体机构政治倾向的标注。
- 提取子主题与命名实体,并通过 Wikidata 进行消歧,以确保语义一致性。
- 通过整合新闻文本、元数据以及来自 Wikidata 的最多两跳邻居,构建特定领域的知识图谱(NeMigKG)。
- 在多个 NeMigKG 变体上使用 TransD 对实体嵌入进行预训练,以评估其对推荐性能的影响。
- 收集了每种语言 3,000 名用户的匿名用户数据,包括人口统计、政治态度、人格特质以及明确的点击反馈。

实验结果
研究问题
- RQ1在新闻文章中引入政治倾向、情感极化信息,如何影响推荐系统性能与多样性?
- RQ2通过 Wikidata 邻居扩展知识图谱,在多大程度上提升了推荐准确率与基于方面(aspect-based)的多样性?
- RQ3具有不同政治取向的媒体机构如何随时间变化报道相同的移民相关实体?
- RQ4用户的社会人口统计与政治数据在多大程度上可用于检测与分析新闻推荐系统中的过滤气泡效应?
- RQ5德国与美国媒体在政治光谱上,其在实体覆盖上的时间趋势如何?
主要发现
- 在 NeMigKG 中加入政治倾向、情感极化与子主题信息,提升了推荐的基于方面的多样性,而对准确率影响不显著。
- 通过 Wikidata 的一跳邻居扩展 NeMigKG 可提升推荐性能,但两跳邻居反而降低性能,表明在小数据集中存在知识过载现象。
- 德语媒体中提及频率最高的 20 个实体主要由中右与中立媒体覆盖,但有显著例外,如安格拉·默克尔与左翼党(Die Linke)在右翼媒体中出现频率更高。
- 在美国媒体中,左倾媒体对最常见实体的报道频率低于中立或右倾媒体。
- 时间趋势分析显示,德国媒体的实体覆盖与重大政治事件(如联邦与欧洲选举,或土耳其在叙利亚北部的军事行动等国际冲突)密切相关。
- 用户人口统计与政治数据的引入,使研究者能够检测到用户特定的偏差模式,支持对个性化推荐对政治极化影响的分析。

更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。