Skip to main content
QUICK REVIEW

[Paper Review] A Longitudinal Measurement Study of 4chan's Politically Incorrect Forum and its Effect on the Web.

Gabriel Emile Hine, Jeremiah Onaolapo|arXiv (Cornell University)|Oct 11, 2016
Hate Speech and Cyberbullying Detection14 references12 citations
TL;DR

This paper presents the first longitudinal measurement study of 4chan's /pol/ forum, analyzing over 8 million posts over 2.5 months to characterize global user distribution, content patterns, and coordinated attacks on social media. It finds that /pol/ drives widespread hate speech and YouTube link sharing, with minimal image repetition, while also demonstrating limited success in poisoning anti-trolling tools through linguistic obfuscation.

ABSTRACT

The discussion board site 4chan has been a part of the dark underbelly of the Internet since its inception, but recent events have brought it to the forefront of the world's collective mind. In particular, /pol/, 4chan's Politically Incorrect board has become a central figure in the outlandish 2016 US election campaign, often linked to the alt-right movement and the rhetoric of hate and racism. Nonetheless, 4chan remains relatively unstudied by the research community. In this paper, we start addressing this gap by analyzing /pol/ along several axes, using a dataset of over 8M posts collected over two and a half months. First, we perform a general characterization, showing that /pol/ users are well distributed around the world and that 4chan's unique features encourage fresh discussions. Then, we analyze content posted on /pol/, finding YouTube links and hate speech to be predominant, that 95\% of images are posted no more than 5 times, and that there are notable differences in the English language used in different parts of the world. Last but not least, we provide quantitative evidence of /pol/'s collective attacks on other social media platforms by analyzing the comments in YouTube videos linked on /pol/. We also present a quantitative case study of /pol/'s attempt to poison anti-trolling tools by altering the language of hate on social media, finding it to be less successful than reported by the popular press. Overall, our analysis not only provides the first measurement study of /pol/, but also insight on online harassment and hate speech trends in online social media.

Motivation & Objective

  • To address the lack of empirical research on /pol/, 4chan's politically incorrect forum, which has gained prominence in online extremism and the 2016 US election discourse.
  • To analyze the global distribution and linguistic diversity of /pol/ users, assessing how 4chan's structural features foster fresh discussions.
  • To investigate the prevalence of hate speech, YouTube link sharing, and image reuse on /pol/ to understand content dynamics.
  • To quantify /pol/'s coordinated attacks on external platforms, particularly YouTube, through comment spamming.
  • To evaluate the effectiveness of linguistic obfuscation in evading anti-trolling tools, challenging popular press claims of success.

Proposed method

  • Collected a longitudinal dataset of over 8 million posts from /pol/ across a 2.5-month period using web scraping and API-based access.
  • Performed geolocation and linguistic analysis on posts to map user distribution and identify regional variations in English usage.
  • Used natural language processing (NLP) techniques to detect hate speech and classify content types, including YouTube link prevalence.
  • Analyzed image frequency distributions to assess reuse patterns, finding 95% of images posted five or fewer times.
  • Tracked comments from YouTube videos linked on /pol/ to measure coordinated downvoting and harassment campaigns.
  • Evaluated linguistic obfuscation strategies by comparing anti-trolling tool detection rates before and after language manipulation on /pol/.

Experimental results

Research questions

  • RQ1How is /pol/ distributed geographically, and what role do 4chan's structural features play in sustaining fresh discussions?
  • RQ2What types of content dominate /pol/, particularly in terms of hate speech, YouTube links, and image sharing?
  • RQ3To what extent does /pol/ coordinate attacks on external platforms like YouTube through comment spamming?
  • RQ4How effective are linguistic obfuscation techniques used on /pol/ in evading automated anti-trolling tools?
  • RQ5How do regional linguistic variations in English on /pol/ reflect global user diversity?

Key findings

  • Users on /pol/ are well-distributed across the globe, indicating a transnational user base rather than a localized phenomenon.
  • Hate speech and YouTube links are the two most predominant content types on /pol/, dominating the forum's discourse.
  • 95% of images posted on /pol/ are shared five or fewer times, suggesting low image reuse and high content freshness.
  • Significant linguistic differences in English usage are observed across geographic regions, reflecting diverse user origins.
  • Coordinated attacks on YouTube comments linked from /pol/ were quantitatively measurable, indicating organized harassment campaigns.
  • Linguistic obfuscation on /pol/ was less effective at evading anti-trolling tools than suggested by media reports, with detection rates remaining high.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.