[Paper Review] Quantifying Search Bias: Investigating Sources of Bias for Political Searches in Social Media
The paper develops a framework to quantify input, ranking, and output bias in social media search (Twitter) for political queries and presents a method to infer tweet/user political bias to assess how data and ranking jointly shape biased results.
Search systems in online social media sites are frequently used to find information about ongoing events and people. For topics with multiple competing perspectives, such as political events or political candidates, bias in the top ranked results significantly shapes public opinion. However, bias does not emerge from an algorithm alone. It is important to distinguish between the bias that arises from the data that serves as the input to the ranking system and the bias that arises from the ranking system itself. In this paper, we propose a framework to quantify these distinct biases and apply this framework to politics-related queries on Twitter. We found that both the input data and the ranking system contribute significantly to produce varying amounts of bias in the search results and in different ways. We discuss the consequences of these biases and possible mechanisms to signal this bias in social media search systems' interfaces.
Motivation & Objective
- Quantify distinct sources of search bias (input, ranking, output) in social media search for political topics.
- Distinguish whether bias originates from data input or the ranking system itself.
- Develop a method to infer political bias of individual Twitter data items (tweets) to support bias quantification.
- Apply the framework to 2016 US political queries on Twitter to measure bias contributions from input data and ranking.
Proposed method
- Propose a three-stage bias quantification framework: input bias, ranking bias, and output bias, based on item-level bias scores.
- Define bias scores for individual data items (tweets) and aggregate them to compute input, output, and ranking biases.
- Use an oracle-like approach where the ranking system is treated as a black box to measure output bias as OB(q,r) and RB(q,r)=OB(q,r)−IB(q).
- Infer political bias of Twitter users (source bias) by computing interest vectors from following patterns and seed Democrat/Republican user sets.
- Compute user bias as Bias(u)=cos_sim(Iu,ID)−cos_sim(Iu,IR) with min–max normalization, and evaluate against human judgments.
Experimental results
Research questions
- RQ1RQ1: How can we quantify the different sources of search engine bias (input, ranking, output)?
- RQ2RQ1b: How biased are political search results on Twitter, and what portion comes from input data vs. the ranking system?
- RQ3RQ2: How can we infer the political bias of individual Twitter items (tweets) to support bias quantification?
Key findings
- Both input data and the ranking system contribute significantly to output bias in Twitter search for political queries.
- The ranking system can shift or alter the polarity of bias relative to the input, varying by candidate and party.
- Differently phrased queries yield significantly different biases, highlighting sensitivity to query formulation.
- The proposed source-bias inference method achieves high coverage and strong correlation with human judgments (AMT) for political bias of users.
- For US senators, the bias inference method achieves 97.96% average coverage and 92.23% average accuracy across groups.
- For self-identified common users, the method achieves 91.12% coverage and 85.73% accuracy on average.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.