[Paper Review] Session-based Cyberbullying Detection in Social Media: A Survey
This survey proposes a session-based cyberbullying detection framework that models repetitive behavior and power imbalance in social media interactions. It reviews data, methods, and best practices for dataset creation, benchmarks state-of-the-art models and large language models on two datasets, and identifies key challenges for future research in session-level cyberbullying detection.
Cyberbullying is a pervasive problem in online social media, where a bully abuses a victim through a social media session. By investigating cyberbullying perpetrated through social media sessions, recent research has looked into mining patterns and features for modeling and understanding the two defining characteristics of cyberbullying: repetitive behavior and power imbalance. In this survey paper, we define the Session-based Cyberbullying Detection framework that encapsulates the different steps and challenges of the problem. Based on this framework, we provide a comprehensive overview of session-based cyberbullying detection in social media, delving into existing efforts from a data and methodological perspective. Our review leads us to propose evidence-based criteria for a set of best practices to create session-based cyberbullying datasets. In addition, we perform benchmark experiments comparing the performance of state-of-the-art session-based cyberbullying detection models as well as large pre-trained language models across two different datasets. Through our review, we also put forth a set of open challenges as future research directions.
Motivation & Objective
- To define a standardized session-based cyberbullying detection framework that captures repetitive behavior and power imbalance in social media interactions.
- To review existing data sources and methodologies for session-based cyberbullying detection across diverse social media platforms.
- To establish evidence-based criteria for creating high-quality, representative session-based cyberbullying datasets.
- To benchmark state-of-the-art session-based models and large pre-trained language models on two public datasets.
- To identify open challenges and future research directions in session-based cyberbullying detection.
Proposed method
- The paper introduces a structured framework that decomposes session-based cyberbullying detection into data collection, feature engineering, modeling, and evaluation stages.
- It analyzes session-level features such as message sequence patterns, linguistic cues, and interaction dynamics to model repetitive behavior and power imbalance.
- The authors propose a set of best practices for dataset construction, including session segmentation, labeling consistency, and data balance across bully-victim dynamics.
- Benchmarking is conducted using state-of-the-art sequential models (e.g., RNNs, Transformers) and large pre-trained language models (e.g., BERT) on two public datasets.
- The evaluation uses standard NLP metrics such as F1-score, precision, and recall to compare model performance across different session configurations.
- The framework integrates insights from computational linguistics, social psychology, and NLP to ensure methodological and ethical robustness.
Experimental results
Research questions
- RQ1How can a standardized framework be designed to model session-based cyberbullying using repetitive behavior and power imbalance?
- RQ2What are the key data and methodological challenges in creating reliable session-based cyberbullying detection datasets?
- RQ3How do state-of-the-art sequential models and large pre-trained language models perform on session-based cyberbullying detection tasks?
- RQ4What evidence-based criteria can guide the creation of high-quality, representative session-based cyberbullying datasets?
- RQ5What are the major open challenges and future research directions in session-based cyberbullying detection?
Key findings
- The proposed session-based detection framework effectively captures the temporal and relational dynamics of cyberbullying through structured session modeling.
- Benchmarking reveals that large pre-trained language models generally outperform traditional sequential models on session-based cyberbullying detection tasks.
- The study identifies significant performance variation across datasets, highlighting the need for standardized evaluation protocols.
- Evidence-based dataset creation criteria improve data quality, consistency, and reproducibility in future research.
- Key challenges remain in handling long sessions, detecting subtle power imbalances, and ensuring model generalization across diverse social media platforms.
- The survey underscores the importance of integrating social psychological insights with NLP techniques for more robust detection systems.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.