[論文レビュー] Comparative Performance of Machine Learning Algorithms in Cyberbullying Detection: Using Turkish Language Preprocessing Techniques
本研究では、ドメイン固有の自然言語処理技術を用いて、トルコ語のソーシャルメディア文書におけるサイバーいじめ検出のための19種類の機械学習アルゴリズムを評価した。Light Gradient Boosting Machine(LGBM)が90.788%の正確性と90.949%のF1スコアを達成し、他のモデルと比較してトルコ語のサイバーいじめ検出において優れた有効性を示した。
With the increasing use of the internet and social media, it is obvious that cyberbullying has become a major problem. The most basic way for protection against the dangerous consequences of cyberbullying is to actively detect and control the contents containing cyberbullying. When we look at today's internet and social media statistics, it is impossible to detect cyberbullying contents only by human power. Effective cyberbullying detection methods are necessary in order to make social media a safe communication space. Current research efforts focus on using machine learning for detecting and eliminating cyberbullying. Although most of the studies have been conducted on English texts for the detection of cyberbullying, there are few studies in Turkish. Limited methods and algorithms were also used in studies conducted on the Turkish language. In addition, the scope and performance of the algorithms used to classify the texts containing cyberbullying is different, and this reveals the importance of using an appropriate algorithm. The aim of this study is to compare the performance of different machine learning algorithms in detecting Turkish messages containing cyberbullying. In this study, nineteen different classification algorithms were used to identify texts containing cyberbullying using Turkish natural language processing techniques. Precision, recall, accuracy and F1 score values were used to evaluate the performance of classifiers. It was determined that the Light Gradient Boosting Model (LGBM) algorithm showed the best performance with 90.788% accuracy and 90.949% F1 Score value.
研究の動機と目的
- トルコ語のソーシャルメディアコンテンツにおけるサイバーいじめ検出に関する包括的な研究の不足に対処すること。
- 多様な機械学習アルゴリズムの性能をトルコ語テキストデータセット上で評価・比較すること。
- NLP前処理を用いたトルコ語のサイバーいじめメッセージを分類するのに最も効果的なアルゴリズムを同定すること。
- トルコ語のデジタル空間におけるより安全なオンラインコミュニケーションを実現するための自動検出システムの改善
提案手法
- 本研究では、トークン化、語形還元、ストップワード除去を含む、トルコ語固有の自然言語処理技術を用いてソーシャルメディア文書を前処理した。
- 洗練されたトルコ語のサイバーいじめデータセットを用いて、19種類の多様な機械学習分類器を訓練および評価した。
- 特徴量抽出にはTF-IDFベクトル化を実施し、テキストをモデル入力用の数値表現に変換した。
- 性能評価には、標準的な指標である適合率、再現率、正確性、F1スコアを用いた。
- 全分類器の性能最適化のため、ハイパーパramータチューニングを適用した。
- 交差検証評価に基づき、LGBMモデルが最良の性能を示したため、採用された。
実験結果
リサーチクエスチョン
- RQ1どの機械学習アルゴリズムがトルコ語のソーシャルメディア文書におけるサイバーいじめ検出で最も優れた性能を示すか?
- RQ2異なるNLP前処理技術が、トルコ語のサイバーいじめ検出モデルの分類性能にどのように影響を与えるか?
- RQ3従来型と木構造ベースの分類器の間で、トルコ語のサイバーいじめコンテンツ検出において、それぞれの有効性はどのように比較できるか?
- RQ4トルコ語のサイバーいじめ検出において、精度、再現率、F1スコアは、モデルごとにどの程度変動するか?
主な発見
- Light Gradient Boosting Machine(LGBM)は、トルコ語のテキストにおけるサイバーいじめ検出で90.788%の最高正確性を達成した。
- LGBMは90.949%の最高F1スコアを記録し、適合率と再現率のバランスが優れていた。
- 他の上位性能を示したモデルにはXGBoostとランダムフォレストがあり、LGBMに比べて競争力はあったが、やや低い性能であった。
- 本研究では、アルゴリズムの選定がトルコ語のサイバーいじめ分類における検出性能に顕著に影響を与えることが確認された。
- トルコ語固有のNLP前処理技術が、モデルの汎化性能と性能向上に不可欠であることが判明した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。