Skip to main content
QUICK REVIEW

[Paper Review] Effects of algorithmic flagging on fairness: quasi-experimental evidence from Wikipedia

Nathan TeBlunthuis, Benjamin Mako Hill|arXiv (Cornell University)|Jun 4, 2020
Wikis in Education and CollaborationSocial Sciences118 references12 citations
TL;DR

This study investigates how algorithmic flagging in Wikipedia moderation affects fairness by reducing reliance on social signals like user registration or profile presence. Using a quasi-experimental regression discontinuity design on RCFilters data, it finds that algorithmic flags increase the likelihood of reverting edits—especially from established editors—while reducing undo rates, suggesting improved fairness for well-intentioned contributors, though effects vary by social signal type.

ABSTRACT

Online community moderators often rely on social signals such as whether or not a user has an account or a profile page as clues that users may cause problems. Reliance on these clues can lead to "overprofiling'' bias when moderators focus on these signals but overlook the misbehavior of others. We propose that algorithmic flagging systems deployed to improve the efficiency of moderation work can also make moderation actions more fair to these users by reducing reliance on social signals and making norm violations by everyone else more visible. We analyze moderator behavior in Wikipedia as mediated by RCFilters, a system which displays social signals and algorithmic flags, and estimate the causal effect of being flagged on moderator actions. We show that algorithmically flagged edits are reverted more often, especially those by established editors with positive social signals, and that flagging decreases the likelihood that moderation actions will be undone. Our results suggest that algorithmic flagging systems can lead to increased fairness in some contexts but that the relationship is complex and contingent.

Motivation & Objective

  • To evaluate whether algorithmic flagging reduces overreliance on social signals in online community moderation.
  • To assess whether algorithmic flagging leads to fairer moderation outcomes by reducing bias against users with positive social signals.
  • To investigate the causal impact of algorithmic flags on moderator actions, particularly regarding edits by established or registered users.
  • To explore how different social signals (e.g., registration status, user pages) interact with algorithmic flagging in shaping moderation fairness.
  • To develop a methodological framework for causal inference in real-world sociotechnical systems without experimental intervention.

Proposed method

  • Utilized a quasi-experimental regression discontinuity design based on arbitrary thresholds in the RCFilters system.
  • Leveraged ORES-generated machine learning scores to identify edits likely to be damaging.
  • Analyzed moderator actions (e.g., reverts, undo actions) on edits just above and below flagging thresholds.
  • Compared moderation outcomes for edits by users with varying social signals (e.g., registered vs. unregistered, with or without user pages).
  • Applied causal inference techniques to estimate the effect of algorithmic flagging on moderation behavior.
  • Used a replication dataset from Harvard Dataverse, including ORES scores, thresholds, and revision data, for reproducibility.

Experimental results

Research questions

  • RQ1Does algorithmic flagging reduce overprofiling by decreasing reliance on social signals like registration status or user pages?
  • RQ2How does algorithmic flagging affect the likelihood of edits being reverted, particularly for established editors?
  • RQ3Does flagging reduce the likelihood that moderation actions are undone, indicating greater decision stability?
  • RQ4How do different types of social signals interact with algorithmic flagging to influence fairness in moderation?
  • RQ5To what extent do flagging thresholds introduce arbitrariness in moderation outcomes?

Key findings

  • Algorithmically flagged edits were significantly more likely to be reverted, especially those by registered editors with positive social signals.
  • Flagged edits by established editors were reverted at a higher rate than unflagged edits, suggesting reduced overreliance on social signals.
  • Moderation actions on flagged edits were less likely to be undone, indicating greater stability and reduced controversy.
  • The fairness impact of flagging varied by social signal: it reduced overprofiling for unregistered users but increased it for users without user pages.
  • The use of arbitrary thresholds in flagging systems led to disproportionate attention on edits just above the threshold, introducing arbitrariness in moderation outcomes.
  • The results suggest that algorithmic flagging can improve fairness in some contexts but may also exacerbate bias depending on the type of social signal involved.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.