[논문 리뷰] Cyberbullying Detection in Social Networks Using Deep Learning Based Models; A Reproducibility Study
이 논문은 위키백과, 트위터, Formspring에서의 딥러닝 기반 사이버 괴롭힘 탐지 결과를 재현하고, 새로운 YouTube 데이터셋(~54,000개의 게시물, ~4,000명의 사용자)으로 평가를 확장하며, 모델의 플랫폼 간 전이(크로스 플랫폼 전이)를 고찰한다.
Cyberbullying is a disturbing online misbehaviour with troubling consequences. It appears in different forms, and in most of the social networks, it is in textual format. Automatic detection of such incidents requires intelligent systems. Most of the existing studies have approached this problem with conventional machine learning models and the majority of the developed models in these studies are adaptable to a single social network at a time. In recent studies, deep learning based models have found their way in the detection of cyberbullying incidents, claiming that they can overcome the limitations of the conventional models, and improve the detection performance. In this paper, we investigate the findings of a recent literature in this regard. We successfully reproduced the findings of this literature and validated their findings using the same datasets, namely Wikipedia, Twitter, and Formspring, used by the authors. Then we expanded our work by applying the developed methods on a new YouTube dataset (~54k posts by ~4k users) and investigated the performance of the models in new social media platforms. We also transferred and evaluated the performance of the models trained on one platform to another platform. Our findings show that the deep learning based models outperform the machine learning models previously applied to the same YouTube dataset. We believe that the deep learning based models can also benefit from integrating other sources of information and looking into the impact of profile information of the users in social networks.
연구 동기 및 목표
- 딥러닝 모델을 이용한 사이버 괴롭힘 탐지에 관한 최근 연구의 발견을 동일한 데이터세트(Wikipedia, Twitter, Formspring)를 사용하여 재현한다.
- 새로운 소셜 플랫폼(YouTube)으로 평가를 확장하고 기존 접근법에 비해 성능을 평가한다.
- 한 플랫폼에서 학습된 모델이 다른 플랫폼으로 얼마나 잘 전이되는지 크로스 플랫폼 전이를 탐구한다.
- 추가 정보(예: 사용자 프로필 데이터)를 포함하는 것이 탐지 성능 향상에 미치는 잠재적 이점을 시사한다.
제안 방법
- 이전 연구에 보고된 딥러닝 기반 사이버 괴롭힘 탐지 방법을 동일한 데이터세트(Wikipedia, Twitter, Formspring)로 재현한다.
- 재현된 방법을 새로운 YouTube 데이터세트(~54,000개의 게시물, ~4,000명의 사용자)에서 적용하고 성능을 평가한다.
- YouTube 데이터세트에서 딥러닝 모델과 전통적인 기계 학습 기준선을 비교한다.
- 한 플랫폼에서 학습된 모델을 다른 플랫폼에서 평가하기 위해 크로스 플랫폼 전이를 조사한다.
- 사용자 프로필 데이터와 같은 추가 정보를 모델 성능에 통합하는 것이 미치는 영향을 논의한다.
실험 결과
연구 질문
- RQ1YouTube 데이터세트에서도 딥러닝 기반 모델이 기존 ML 모델보다 우수한가?
- RQ2Wikipedia, Twitter, Formspring의 발견을 동일한 데이터세트로 재현할 수 있는가?
- RQ3한 소셜 플랫폼에서 학습된 모델이 다른 플랫폼으로 얼마나 잘 전이되는가?
- RQ4사용자 프로필 정보를 포함하는 것이 사이버 괴롭힘 탐지 성능에 미치는 잠재적 영향은 무엇인가?
주요 결과
- DL 기반 모델이 YouTube 데이터세트에 대해 이전에 적용된 머신러닝 모델보다 우수한 성능을 보인다.
- 저자들은 동일한 데이터세트를 사용하여 Wikipedia, Twitter, Formspring의 연구 결과를 성공적으로 재현했다.
- 새로운 YouTube 데이터세트에 적용했을 때 DL 모델은 전통적인 ML 기준선에 비해 우수한 성능을 나타냈다.
- 한 플랫폼에서 학습된 모델을 다른 플랫폼으로 전이하여 평가할 수 있다.
- 사용자 프로필 정보와 같은 추가 정보를 통합하는 것이 모델 성능에 이익을 줄 수 있다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.