Skip to main content
QUICK REVIEW

[论文解读] Quantifying Search Bias: Investigating Sources of Bias for Political Searches in Social Media

Juhi Kulshrestha, Motahhare Eslami|arXiv (Cornell University)|Apr 5, 2017
Complex Network Analysis Techniques参考文献 43被引用 134
一句话总结

本文开发了一个框架,用于量化社交媒体搜索(Twitter)在政治性查询中的输入、排序和输出偏差,并提出一种推断 tweet/user 政治偏向以评估数据和排序如何共同形塑偏见结果。

ABSTRACT

Search systems in online social media sites are frequently used to find information about ongoing events and people. For topics with multiple competing perspectives, such as political events or political candidates, bias in the top ranked results significantly shapes public opinion. However, bias does not emerge from an algorithm alone. It is important to distinguish between the bias that arises from the data that serves as the input to the ranking system and the bias that arises from the ranking system itself. In this paper, we propose a framework to quantify these distinct biases and apply this framework to politics-related queries on Twitter. We found that both the input data and the ranking system contribute significantly to produce varying amounts of bias in the search results and in different ways. We discuss the consequences of these biases and possible mechanisms to signal this bias in social media search systems' interfaces.

研究动机与目标

  • 量化社交媒体搜索中政治话题的不同来源偏差(输入、排序、输出)。
  • 区分偏差是源自数据输入还是排序系统本身。
  • 开发一种推断单个 Twitter 数据项(推文)政治偏向以支持偏差量化的方法。
  • 将该框架应用于 Twitter 上的2016年美国政治查询,以衡量来自输入数据和排序的偏见贡献。

提出的方法

  • 提出一个三阶段偏差量化框架:输入偏差、排序偏差和输出偏差,基于项级偏差分数。
  • 为单个数据项(推文)定义偏差分数并聚合它们以计算输入、输出和排序偏差。
  • 采用类 oracle 的方法,将排序系统视为黑箱,以将输出偏差测量为 OB(q,r) 且 RB(q,r)=OB(q,r)−IB(q)。
  • 通过根据关注模式计算兴趣向量并使用种子民主党/共和党用户集合来推断 Twitter 用户的政治偏向(源偏差)。
  • 在最小-最大归一化下,将用户偏置计算为 Bias(u)=cos_sim(Iu,ID)−cos_sim(Iu,IR),并与人类判断进行评估。

实验结果

研究问题

  • RQ1RQ1:我们如何量化搜索引擎偏差的不同来源(输入、排序、输出)?
  • RQ2RQ1b:Twitter 的政治搜索结果有多偏?其中有多少来自输入数据 vs. 排序系统?
  • RQ3RQ2:我们如何推断单个 Twitter 项(推文)的政治偏向以支持偏差量化?

主要发现

  • 输入数据和排序系统都对政治查询的 Twitter 搜索中的输出偏差有显著贡献。
  • 排序系统相对于输入可以改变偏见的极性,且因候选人和政党而异。
  • 不同措辞的查询产生显著不同的偏差,突出对查询表述的敏感性。
  • 所提出的源偏差推断方法在用户的政治偏见方面实现了高覆盖率并与人类判决(AMT)有较强相关性。
  • 对于美国参议员,该偏见推断方法在各组中实现了97.96%的平均覆盖率和92.23%的平均准确性。
  • 对于自我认定的普通用户,该方法平均覆盖率达到91.12%、准确性为85.73%。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。