Skip to main content
QUICK REVIEW

[论文解读] A Weakly-Labeled Stance Dataset during the 2019 South American Protests

Ramón Villa-Cox, Helen Zeng|arXiv (Cornell University)|Apr 5, 2021
Social Media and Politics参考文献 19被引用 5
一句话总结

本文利用2019年南美洲抗议期间的Twitter数据,通过用户对政治人物的支持和话题标签活动,构建了一个弱标签化的立场数据集。采用基于机器翻译的方法检测语言极化,并通过用户在媒体集群间的移动性分析新闻消费模式,发现各国普遍存在强烈的过滤气泡效应和意识形态隔离,其中玻利维亚的极化程度最高,智利最低。

ABSTRACT

Research across different disciplines has documented the expanding polarization in social media. However, much of it focused on the US political system or its culturally controversial topics. In this work, we explore polarization on Twitter in a different context, namely the protest that paralyzed several countries in the South American region in 2019. By leveraging users' endorsement of politicians' tweets and hashtag campaigns with defined stances towards the protest (for or against), we construct a weakly labeled stance dataset with millions of users. We explore polarization in two related dimensions: language and news consumption patterns. In terms of linguistic polarization, we apply recent insights that leveraged machine translation methods, showing that the two communities speak consistently "different" languages, mainly along ideological lines (e.g., fascist translates to communist). Our results indicate that this recently-proposed methodology is also informative in different languages and contexts than originally applied. In terms of news consumption patterns, we cluster news agencies based on homogeneity of their user bases and quantify the observed polarization in its consumption. We find empirical evidence of the "filter bubble" phenomenon during the event, as we not only show that the user bases are homogeneous in terms of stance, but the probability that a user transitions from media of different clusters is low.

研究动机与目标

  • 利用Twitter数据研究2019年南美洲抗议期间用户行为的极化现象。
  • 通过在用户生成内容上应用基于机器翻译的方法,探索不同意识形态立场下的语言极化。
  • 通过基于用户立场同质性的聚类方法,分析新闻消费模式中的极化现象。
  • 通过用户在媒体集群间的移动性评估过滤气泡的存在。
  • 考察区域性和国际媒体机构(如RT en Español和TeleSUR)在塑造信息生态系统中的作用。

提出的方法

  • 根据用户转发政治人物推文及使用话题标签的情况,将用户分类为亲政府或反政府立场。
  • 分别在每类立场群体的语料上训练词嵌入,以捕捉语言差异。
  • 在两个嵌入空间之间学习翻译矩阵,以识别系统性语义偏移(例如,'fascist'被翻译为'communist')。
  • 基于用户群体在立场上的同质性,对新闻媒体进行聚类。
  • 构建转移矩阵以建模用户在媒体集群间的移动性,通过移动性指数(IR、ML、MR)量化过滤气泡效应。
  • 该方法在玻利维亚、智利、哥伦比亚和厄瓜多尔四个国家中应用,以比较极化水平。

实验结果

研究问题

  • RQ1在2019年南美洲抗议期间,持对立立场的用户在语言使用上有多大的差异?
  • RQ2基于机器翻译的方法如何揭示用户语言中的意识形态语义偏移?
  • RQ3用户新闻消费模式的极化程度如何?是否存在过滤气泡的证据?
  • RQ4区域性与国际媒体机构(如RT en Español和TeleSUR)在信息传播与用户参与中发挥何种作用?
  • RQ5玻利维亚、智利、哥伦比亚和厄瓜多尔在语言和新闻消费两个维度上的极化水平如何变化?

主要发现

  • 语言极化现象明显,存在系统性语义偏移,例如对立立场中'fascist'被翻译为'communist','police'被翻译为'vandals'。
  • 基于机器翻译的方法成功检测出西班牙语中的意识形态语言差异,拓展了其在非英语语境下的适用性。
  • 媒体集群沿意识形态和地理边界形成,俄罗斯和委内瑞拉媒体(如RT en Español、TeleSUR)构成一个左倾集群。
  • 总体用户移动性较低,静止比率(IR)从智利的72.74%到玻利维亚的95.68%不等,表明存在强烈的过滤气泡效应。
  • 用户从右倾媒体转向左倾媒体的概率高于反向,ML > MR在玻利维亚、智利、哥伦比亚和厄瓜多尔均成立。
  • 极化水平在玻利维亚最高,智利最低,哥伦比亚和厄瓜多尔处于相似的中等水平,且在语言和媒体消费两个维度上均保持一致。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。