[논문 리뷰] Comparative Performance of Machine Learning Algorithms in Cyberbullying Detection: Using Turkish Language Preprocessing Techniques
이 연구는 도메인 특화 자연어 처리 기법을 사용하여 터키어 소셜미디어 텍스트에서 사이버불링 감지에 대해 19개의 기계학습 알고리즘을 평가한다. 라이트 그래디언트 부스팅 머신(LGBM)이 90.788%의 정확도와 90.949%의 F1 스코어로 가장 높은 성능을 기록하여, 다른 모델들과 비교해 터키어 사이버불링 감지에 효과적임을 입증한다.
With the increasing use of the internet and social media, it is obvious that cyberbullying has become a major problem. The most basic way for protection against the dangerous consequences of cyberbullying is to actively detect and control the contents containing cyberbullying. When we look at today's internet and social media statistics, it is impossible to detect cyberbullying contents only by human power. Effective cyberbullying detection methods are necessary in order to make social media a safe communication space. Current research efforts focus on using machine learning for detecting and eliminating cyberbullying. Although most of the studies have been conducted on English texts for the detection of cyberbullying, there are few studies in Turkish. Limited methods and algorithms were also used in studies conducted on the Turkish language. In addition, the scope and performance of the algorithms used to classify the texts containing cyberbullying is different, and this reveals the importance of using an appropriate algorithm. The aim of this study is to compare the performance of different machine learning algorithms in detecting Turkish messages containing cyberbullying. In this study, nineteen different classification algorithms were used to identify texts containing cyberbullying using Turkish natural language processing techniques. Precision, recall, accuracy and F1 score values were used to evaluate the performance of classifiers. It was determined that the Light Gradient Boosting Model (LGBM) algorithm showed the best performance with 90.788% accuracy and 90.949% F1 Score value.
연구 동기 및 목표
- 터키어 소셜미디어 콘텐츠에서 사이버불링 감지에 대한 종합적인 연구 부족을 해결하기 위해.
- 다양한 기계학습 알고리즘의 성능을 터키어 텍스트 데이터셋에서 평가하고 비교하기 위해.
- NLP 전처리 기법을 사용해 터키어에서 사이버불링 메시지를 분류하는 데 가장 효과적인 알고리즘을 특정하기 위해.
- 터키 디지털 환경에서 더 안전한 온라인 소통을 위한 자동화된 감지 시스템을 향상시키기 위해.
제안 방법
- 연구는 터키어 전용 자연어 처리 기법을 적용해 소셜미디어 텍스트를 전처리하였으며, 토큰화, 어간 추출, 불용어 제거를 포함한다.
- 정제된 터키어 사이버불링 데이터셋에 대해 다양한 19개의 기계학습 분류기 모델을 훈련하고 평가하였다.
- 모델 입력을 위해 텍스트를 수치적 표현으로 변환하기 위해 TF-IDF 벡터화를 사용해 특징 추출을 수행하였다.
- 표준 평가 지표인 정밀도, 재현율, 정확도, F1 스코어를 사용해 성능을 측정하였다.
- 모든 분류기에서 성능 최적화를 위해 초모수 튜닝을 적용하였다.
- 교차 검증 평가 기반으로 LGBM 모델이 최고의 성능을 기록하여 최종 선택되었다.
실험 결과
연구 질문
- RQ1어느 기계학습 알고리즘이 터키어 소셜미디어 텍스트에서 사이버불링 감지에 가장 잘 작동하는가?
- RQ2다양한 NLP 전처리 기법이 터키어 사이버불링 감지 모델의 분류 성능에 어떤 영향을 미치는가?
- RQ3기존 기반 및 트리 기반 분류기의 상대적 효과성은 터키어 사이버불링 콘텐츠 감지에서 어떻게 나타나는가?
- RQ4터키어 사이버불링 감지 시 모델 간 정밀도, 재현율, F1 스코어는 얼마나 다를까?
주요 결과
- 라이트 그래디언트 부스팅 머신(LGBM)이 터키어 텍스트에서 사이버불링 감지에 90.788%의 최고 정확도를 기록하였다.
- LGBM는 정밀도와 재현율의 균형이 우수한 90.949%의 최고 F1 스코어를 기록하였다.
- 다른 상위 성능 모델로는 XGBoost와 랜덤 포레스트가 있었으며, LGBM에 비해 경쟁력은 있었지만 성능이 낮았다.
- 연구는 알고리즘 선택이 터키어 사이버불링 분류에서 감지 성능에 상당한 영향을 미친다는 것을 확인하였다.
- 터키어 전용 NLP 전처리 기법이 모델의 일반화 능력과 성능 향상에 필수적이었다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.