Skip to main content
QUICK REVIEW

[Paper Review] Large Scale Crowdsourcing and Characterization of Twitter Abusive Behavior

Antigoni-Maria Founta, Constantinos Djouvas|arXiv (Cornell University)|Feb 1, 2018
Hate Speech and Cyberbullying Detection4 citations
TL;DR

This paper presents a large-scale, crowdsourced methodology for labeling abusive behavior on Twitter, using iterative annotation and boosted sampling to improve detection of rare abusive categories. It produces a robust, publicly available dataset of 80,000 tweets labeled with refined categories—primarily 'abusive' and 'hateful'—after merging and eliminating overlapping labels through statistical analysis of annotator agreement.

ABSTRACT

In recent years, offensive, abusive and hateful language, sexism, racism and other types of aggressive and cyberbullying behavior have been manifesting with increased frequency, and in many online social media platforms. In fact, past scientific work focused on studying these forms in popular media, such as Facebook and Twitter. Building on such work, we present an 8-month study of the various forms of abusive behavior on Twitter, in a holistic fashion. Departing from past work, we examine a wide variety of labeling schemes, which cover different forms of abusive behavior, at the same time. We propose an incremental and iterative methodology, that utilizes the power of crowdsourcing to annotate a large scale collection of tweets with a set of abuse-related labels. In fact, by applying our methodology including statistical analysis for label merging or elimination, we identify a reduced but robust set of labels. Finally, we offer a first overview and findings of our collected and annotated dataset of 100 thousand tweets, which we make publicly available for further scientific exploration.

Motivation & Objective

  • To address the challenge of inconsistent labeling of abusive behavior on social media, particularly the difficulty in distinguishing nuanced categories like hate speech, cyberbullying, and offensive language.
  • To develop a scalable, cost-effective crowdsourcing methodology that maintains annotation quality while ensuring sufficient representation of low-frequency abusive categories.
  • To create a large, publicly available, and statistically validated dataset of Twitter tweets annotated for abusive behavior to support future research in automated detection systems.
  • To reduce label ambiguity by identifying and merging overlapping or redundant labels through statistical analysis of annotator agreement.

Proposed method

  • An incremental and iterative crowdsourcing framework was designed, involving multiple rounds of labeling with feedback to improve consistency among annotators.
  • A boosted sampling strategy was employed to increase the proportion of rare abusive categories (e.g., hateful, abusive) in the dataset, countering the low natural occurrence of such content.
  • A custom-built data collection platform was developed to optimize cost, time, and annotation quality by managing worker assignment, payment, and task routing.
  • Label merging and elimination were performed using statistical analysis of inter-annotator agreement, including overwhelming (≥4/5), strong (≥3/5), and simple (≥2/5) agreement scores.
  • The final labeling scheme was reduced to a minimal yet robust set of labels—'abusive', 'hateful', 'spam', 'normal'—after eliminating less distinct or redundant labels like 'cyberbullying'.
  • The dataset was constructed from two subsets: 70,000 random tweets and 10,000 boosted samples, with the latter enriched for abusive content to ensure statistical power for minority classes.

Experimental results

Research questions

  • RQ1How can crowdsourced labeling be optimized to reduce confusion among annotators when distinguishing between nuanced forms of abusive language on Twitter?
  • RQ2What sampling strategy ensures sufficient representation of low-frequency abusive categories without introducing bias in the final dataset?
  • RQ3Which labels for abusive behavior are statistically distinct and reliable, and which can be merged or eliminated based on inter-annotator agreement?
  • RQ4To what extent does a boosted sampling approach improve the detectability and characterization of abusive content compared to random sampling?
  • RQ5How can a large-scale, high-quality, and publicly available dataset of abusive Twitter content be systematically constructed and validated?

Key findings

  • Over 55% of the 80,000 tweets achieved overwhelming agreement (≥4 out of 5 annotators) on their label, indicating strong inter-annotator reliability.
  • The final dataset contains 59% normal, 22.5% spam, 11% abusive, and 7.5% hateful tweets, with the remaining 20% labeled as inappropriate or ambiguous.
  • The boosted sample subset (10,000 tweets) contained nearly half (48%) of all inappropriate content, compared to only 4% in the random sample, demonstrating the effectiveness of the sampling strategy.
  • After statistical analysis, labels such as 'cyberbullying' were eliminated due to low distinctiveness and high overlap with 'abusive' and 'hateful' labels.
  • The final annotation scheme was reduced to four core labels: 'abusive', 'hateful', 'spam', and 'normal', which were found to be the most reliable and representative for characterizing abusive behavior.
  • The dataset is publicly available at http://ow.ly/BqCf30jqffN, and the platform code is open-sourced at http://ow.ly/TnuU30jqf7g for reuse by the research community.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.