Skip to main content
QUICK REVIEW

[Paper Review] Understanding the (In)Effectiveness of Content Moderation: A Case Study of Facebook in the Context of the U.S. Capitol Riot

Ian Goldstein, Laura Edelson|arXiv (Cornell University)|Jan 6, 2023
Hate Speech and Cyberbullying Detection4 citations
TL;DR

This study develops a novel method to infer Facebook content removal timing and impact using public engagement data from 2.5K U.S. news sources around the January 6 Capitol Riot. It finds that even rapid moderation prevents only 21% of predicted engagement, as most virality occurs within 30 hours—highlighting systemic limits of content moderation in crises and urging algorithmic changes to slow content diffusion.

ABSTRACT

Social media networks commonly employ content moderation as a tool to limit the spread of harmful content. However, the efficacy of this strategy in limiting the delivery of harmful content to users is not well understood. In this paper, we create a framework to quantify the efficacy of content moderation and use our metrics to analyze content removal on Facebook within the U.S. news ecosystem. In a data set of over 2M posts with 1.6B user engagements collected from 2,551 U.S. news sources before and during the Capitol Riot on January 6, 2021, we identify 10,811 removed posts. We find that the active engagement life cycle of Facebook posts is very short, with 90% of all engagement occurring within the first 30 hours after posting. Thus, even relatively quick intervention allowed significant accrual of engagement before removal, and prevented only 21% of the predicted engagement potential during a baseline period before the U.S. Capitol attack. Nearly a week after the attack, Facebook began removing older content, but these removals occurred so late in these posts' engagement life cycles that they disrupted less than 1% of predicted future engagement, highlighting the limited impact of this intervention. Content moderation likely has limits in its ability to prevent engagement, especially in a crisis, and we recommend that other approaches such as slowing down the rate of content diffusion be investigated.

Motivation & Objective

  • To assess the real-world effectiveness of Facebook's content moderation in limiting harmful content exposure during crises.
  • To develop a method for inferring post removal timing and engagement impact using publicly available engagement data.
  • To quantify how much engagement occurs before removal and how much is prevented, using predictive models of post virality.
  • To evaluate whether content moderation can meaningfully disrupt harmful content spread in high-engagement events like the Capitol Riot.
  • To recommend improved transparency and alternative strategies, such as slowing content diffusion, to enhance platform safety.

Proposed method

  • Inferred content removals by detecting post disappearance from daily snapshots of active Facebook posts from 2,551 U.S. news sources.
  • Developed a predictive model for post engagement potential using historical engagement patterns from the same page, with 3.2–4.5% error for normal and viral posts.
  • Defined two key metrics: 'accrued engagement' (engagement before removal) and 'prevented engagement' (predicted potential minus actual engagement).
  • Used a baseline period to estimate normal engagement life cycles, finding 90% of engagement occurs within 30 hours of posting.
  • Applied the same metrics to the crisis period, comparing removal timing and impact before and after Facebook’s policy changes on January 12.
  • Classified sources by reputation for factualness using third-party data to assess whether high-risk content was disproportionately removed.

Experimental results

Research questions

  • RQ1How quickly does Facebook remove harmful content from U.S. news publishers and influencers, and how much engagement occurs before removal?
  • RQ2To what extent does content moderation prevent exposure to harmful content, measured by predicted versus actual engagement?
  • RQ3How does the effectiveness of content moderation change during a high-engagement crisis like the January 6 Capitol Riot?
  • RQ4How late are removals occurring relative to the peak of a post’s engagement life cycle, and what is their actual impact?
  • RQ5What metrics and transparency practices could improve the evaluation and design of content moderation systems?

Key findings

  • The active engagement life cycle of Facebook posts is extremely short, with 90% of all engagement occurring within the first 30 hours after posting.
  • Despite a median removal time of 21 hours, content moderation prevented only 21.2% of the predicted engagement potential during normal periods.
  • During the Capitol Riot crisis, removals were initially slow and ineffective, with only 21% of predicted engagement prevented, similar to baseline levels.
  • Facebook began removing older content six days after the attack, but these late removals disrupted less than 1% of predicted future engagement.
  • Even with policy changes and announcements, content moderation failed to keep pace with the speed of content diffusion during the crisis.
  • The results suggest that content moderation alone is insufficient to prevent harmful content spread and that alternative strategies—such as slowing content diffusion—should be prioritized.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.