Skip to main content
QUICK REVIEW

[论文解读] The drivers of online polarization: fitting models to data

Carlo Michele Valensise, Matteo Cinelli|arXiv (Cornell University)|May 31, 2022
Opinion Dynamics and Social Influence被引用 5
一句话总结

本文提出了一种灵活的意见动态模型,整合了人类认知偏见与算法推荐机制,以模拟在线极化现象。通过使用 Jensen-Shannon 散度,定量比较模拟意见分布与来自四个社交媒体平台的实证数据,研究识别出关键驱动因素——尤其是控制极化水平的算法参数 γ——并提出了一套数据驱动的框架,用于优化推荐系统以减少极端极化。

ABSTRACT

Users online tend to join polarized groups of like-minded peers around shared narratives, forming echo chambers. The echo chamber effect and opinion polarization may be driven by several factors including human biases in information consumption and personalized recommendations produced by feed algorithms. Until now, studies have mainly used opinion dynamic models to explore the mechanisms behind the emergence of polarization and echo chambers. The objective was to determine the key factors contributing to these phenomena and identify their interplay. However, the validation of model predictions with empirical data still displays two main drawbacks: lack of systematicity and qualitative analysis. In our work, we bridge this gap by providing a method to numerically compare the opinion distributions obtained from simulations with those measured on social media. To validate this procedure, we develop an opinion dynamic model that takes into account the interplay between human and algorithmic factors. We subject our model to empirical testing with data from diverse social media platforms and benchmark it against two state-of-the-art models. To further enhance our understanding of social media platforms, we provide a synthetic description of their characteristics in terms of the model's parameter space. This representation has the potential to facilitate the refinement of feed algorithms, thus mitigating the detrimental effects of extreme polarization on online discourse.

研究动机与目标

  • 开发一个统一的意见动态模型,以同时捕捉人类行为偏见与在线信息消费中的算法影响。
  • 解决现有意见动态模型在真实社交媒体数据上缺乏系统性、定量验证的问题。
  • 通过将模型输出与多样化平台的实证意见分布进行比较,识别出驱动极化的最关键参数。
  • 基于模型拟合结果,对社交媒体平台进行参数化表征,从而支持有针对性的算法干预。
  • 证明通过调节算法影响参数 γ 可有效降低极化水平,为缓解有害的在线动态提供可行路径。

提出的方法

  • 构建一个灵活的意见动态模型,整合同质性偏好、社会传播效应与算法偏见,其中关键参数 γ 控制同质化邻居的影响程度。
  • 采用 Jensen-Shannon (JS) 散度,定量测量模拟意见分布与来自四个平台(Facebook、Twitter、Reddit 和 Gab)的实证数据之间的相似性。
  • 在模型参数空间中进行网格搜索,以识别在各平台上最能复现观测到的极化模式的最优参数配置。
  • 与两项最先进的模型(Arruda 等 [22] 及其他模型)进行对比,验证本模型在捕捉极端极化方面的优越性能。
  • 利用合成参数空间表示,以意见动态机制为基准,描述各平台的特性。
  • 将该方法应用于社区级与主题特定的数据,确保在不同数据采集层级下的稳健性。

实验结果

研究问题

  • RQ1哪些人类与算法因素的组合最能解释多样化社交媒体平台上在线极化的产生?
  • RQ2统一的意见动态模型在多大程度上能复现社交媒体上实测的真实意见分布?
  • RQ3控制同质化同伴相对影响力的算法参数 γ 在塑造极化水平方面发挥何种作用?
  • RQ4与现有最先进的模型相比,本模型在捕捉实证极化模式方面的表现如何?
  • RQ5能否建立一种定量、数据驱动的方法,将社交媒体平台映射到意见动态模型的参数空间中,以获得可操作的洞见?

主要发现

  • 所提出的模型在复现所有四个测试平台(Facebook、Twitter、Reddit、Gab)的实证意见分布方面,显著优于现有模型。
  • 控制算法强化同质化影响的参数 γ 被识别为降低极化的最关键因素,其最优值可使 JS 散度最小化。
  • 在 Reddit 和 Twitter 上,帖子分布函数的影响较小,而前 5% 最优配置中的相位分布 φ 展现出强烈局域化特征,表明其重要性。
  • 重连机制与用户发帖行为(pol、uni、sim 类型)对于准确复现意见分布至关重要,尤其在 Reddit 上表现显著。
  • 该模型在不同数据采集层级(社区级,如 Gab 和 Reddit;主题特定,如 Twitter 和 Facebook)上均表现出稳健的数据拟合能力,证明了其泛化能力。
  • 并不存在一个能捕捉所有现实世界极化动态的“万能模型”,但本方法支持基于实证验证的系统性模型比较与优化。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。