Skip to main content
QUICK REVIEW

[Paper Review] The paradigm-shift of social spambots: Evidence, theories, and tools for the arms race

Stefano Cresci, Roberto Di Pietro|Technical University of Denmark, DTU Orbit (Technical University of Denmark, DTU)|Jan 11, 2017
Spam and Phishing DetectionComputer Science47 references227 citations
TL;DR

This paper provides empirical evidence of a new wave of social spambots on Twitter that evade both platform and human detection, evaluates existing detection tools, analyzes crowdsourced human performance, and advocates for group-behavior based annotation and new detection approaches.

ABSTRACT

Recent studies in social media spam and automation provide anecdotal argumentation of the rise of a new generation of spambots, so-called social spambots. Here, for the first time, we extensively study this novel phenomenon on Twitter and we provide quantitative evidence that a paradigm-shift exists in spambot design. First, we measure current Twitter's capabilities of detecting the new social spambots. Later, we assess the human performance in discriminating between genuine accounts, social spambots, and traditional spambots. Then, we benchmark several state-of-the-art techniques proposed by the academic literature. Results show that neither Twitter, nor humans, nor cutting-edge applications are currently capable of accurately detecting the new social spambots. Our results call for new approaches capable of turning the tide in the fight against this raising phenomenon. We conclude by reviewing the latest literature on spambots detection and we highlight an emerging common research trend based on the analysis of collective behaviors. Insights derived from both our extensive experimental campaign and survey shed light on the most promising directions of research and lay the foundations for the arms race against the novel social spambots. Finally, to foster research on this novel phenomenon, we make publicly available to the scientific community all the datasets used in this study.

Motivation & Objective

  • Demonstrate the existence of a novel wave of social spambots and their detection challenges.
  • Evaluate current Twitter and human capabilities in detecting social spambots.
  • Review and critique traditional detection tools and features versus new, group-based approaches.
  • Provide datasets and guidelines to advance annotation and benchmarking in spambots research.
  • Lay foundations for ongoing arms race strategies against evolving social spambots.

Proposed method

  • Construct and analyze multiple Twitter datasets including genuine accounts, traditional spambots, and novel social spambots across three groups.
  • Assess Twitter’s ability to suspend malicious accounts using API response codes to estimate survivability.
  • Conduct a crowdsourcing detection campaign with trusted contributors to classify accounts and measure human performance.
  • Benchmark established spambot detection techniques (BotOrNot?, Yang et al. classifier, and unsupervised/graph-based methods) on the new spambots datasets.
  • Propose and implement alternative annotation methodologies focusing on group behavior similarities to improve ground-truth datasets.
Figure 1: Survival rates for different types of accounts.
Figure 1: Survival rates for different types of accounts.

Experimental results

Research questions

  • RQ1RQ1: To what extent is Twitter capable of detecting and removing social spambots?
  • RQ2RQ2: Do humans succeed in detecting social spambots in the wild?
  • RQ3RQ3: Can humans distinguish traditional spambots, social spambots, and genuine accounts?
  • RQ4RQ4: Are state-of-the-art detection tools able to detect social spambots?
  • RQ5RQ5: What emerging methodological directions can effectively counter social spambots?

Key findings

  • Genuine accounts show high survival on Twitter (96.5%), while fake followers and some traditional spambots are largely detected or have high suspension rates.
  • Social spambots have survival rates similar to genuine accounts (95.2%–99.6%), indicating greater evasion of platform detection than traditional spambots.
  • Crowdworkers achieve high accuracy on traditional spambots (≈0.91–0.92) and genuine accounts (≈0.92) but perform poorly on social spambots (≈0.24), with low inter-rater agreement for social spambots (κ ≈ 0.186).
  • Established tools show limited success against social spambots; BotOrNot? and Yang et al. classifier perform poorly on social spambots, especially recall.
  • An unsupervised graph-clustering approach (fastgreedy) to account-feature graphs achieves strong detection (MCC ≈ 0.886 for test set #1, ≈0.847 for test set #2), outperforming several supervised and text-heavy methods.
  • The paper advocates group-behavior based annotation and provides publicly released annotated datasets to support such approaches.
Figure 2: Dataset composition for the crowdsourcing experiment.
Figure 2: Dataset composition for the crowdsourcing experiment.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.