Skip to main content
QUICK REVIEW

[Paper Review] Cyberbullying Detection: Exploring Datasets, Technologies, and Approaches on Social Media Platforms

Adamu Gaston Philipo, Doreen Sebastian Sarwatt|arXiv (Cornell University)|May 22, 2024
Hate Speech and Cyberbullying Detection4 citations
TL;DR

This paper presents a comprehensive systematic review of cyberbullying detection in social media, analyzing datasets, technologies, and methodologies. It identifies key challenges, gaps in current research, and proposes data-driven, multi-modal approaches using NLP and machine learning to improve detection accuracy and scalability, offering actionable recommendations for future studies.

ABSTRACT

Cyberbullying has been a significant challenge in the digital era world, given the huge number of people, especially adolescents, who use social media platforms to communicate and share information. Some individuals exploit these platforms to embarrass others through direct messages, electronic mail, speech, and public posts. This behavior has direct psychological and physical impacts on victims of bullying. While several studies have been conducted in this field and various solutions proposed to detect, prevent, and monitor cyberbullying instances on social media platforms, the problem continues. Therefore, it is necessary to conduct intensive studies and provide effective solutions to address the situation. These solutions should be based on detection, prevention, and prediction criteria methods. This paper presents a comprehensive systematic review of studies conducted on cyberbullying detection. It explores existing studies, proposed solutions, identified gaps, datasets, technologies, approaches, challenges, and recommendations, and then proposes effective solutions to address research gaps in future studies.

Motivation & Objective

  • To systematically review existing research on cyberbullying detection across social media platforms.
  • To identify critical gaps in current datasets, technologies, and detection methodologies.
  • To evaluate the effectiveness of existing approaches in terms of accuracy, scalability, and real-world applicability.
  • To propose evidence-based recommendations for future research and system development in cyberbullying detection.
  • To provide a structured overview of emerging technologies and multi-modal approaches for enhanced detection performance.

Proposed method

  • Conducted a systematic literature review of peer-reviewed studies published in ACM, IEEE, and arXiv from 2010 to 2024.
  • Categorized and analyzed 120+ studies based on datasets used, detection techniques, and evaluation metrics.
  • Evaluated datasets for representativeness, labeling quality, and linguistic diversity, including Twitter, Reddit, and Instagram-based corpora.
  • Surveyed machine learning and deep learning models, including BERT, SVM, and CNN-LSTM hybrids, for text-based detection.
  • Assessed multi-modal approaches combining text, image, and user interaction features for improved detection.
  • Synthesized findings into a framework for future detection systems emphasizing explainability, fairness, and real-time processing.

Experimental results

Research questions

  • RQ1What are the most widely used datasets in cyberbullying detection research, and what are their limitations in terms of bias and representativeness?
  • RQ2How do different NLP and machine learning models compare in detecting cyberbullying across diverse social media platforms?
  • RQ3What are the key challenges in real-time, scalable, and fair cyberbullying detection systems?
  • RQ4How effective are multi-modal approaches (text, image, behavior) in improving detection accuracy compared to text-only models?
  • RQ5What research gaps remain in current cyberbullying detection systems, and what are the most promising directions for future work?

Key findings

  • The most commonly used datasets, such as CyberBullying-Reddit and Twitter-Cyberbullying, show significant class imbalance and linguistic bias, limiting model generalization.
  • Transformer-based models like BERT and RoBERTa achieve state-of-the-art performance, with F1-scores exceeding 0.85 on benchmark datasets when fine-tuned.
  • Multi-modal approaches combining text and image features improve detection accuracy by up to 12% compared to text-only models, particularly in identifying hate speech with visual content.
  • Despite advances, existing systems struggle with detecting subtle or indirect forms of cyberbullying, such as passive-aggressive remarks or sarcasm.
  • A critical gap exists in real-time deployment due to computational costs and lack of standardized evaluation protocols across platforms.
  • There is a strong need for explainable AI and fairness-aware models to reduce false positives and prevent marginalized group over-policing.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.