[論文レビュー] Text Classification Using Hybrid Machine Learning Algorithms on Big Data
本論文は、テキストマイニング技術を組み合わせたハイブリッド機械学習モデルを提案し、ビッグデータにおけるテキスト分類の正確性と効率性を向上させることを目的としている。この手法は計算複雑性を低減し、WEKAとJavaを用いたソーシャルメディアのテキストデータセット上で、96.76%の正確性を達成した。これは、単独のナイーブベイズ(61.45%)やSVM(69.21%)を著しく上回っている。
Recently, there are unprecedented data growth originating from different online platforms which contribute to big data in terms of volume, velocity, variety and veracity (4Vs). Given this nature of big data which is unstructured, performing analytics to extract meaningful information is currently a great challenge to big data analytics. Collecting and analyzing unstructured textual data allows decision makers to study the escalation of comments/posts on our social media platforms. Hence, there is need for automatic big data analysis to overcome the noise and the non-reliability of these unstructured dataset from the digital media platforms. However, current machine learning algorithms used are performance driven focusing on the classification/prediction accuracy based on known properties learned from the training samples. With the learning task in a large dataset, most machine learning models are known to require high computational cost which eventually leads to computational complexity. In this work, two supervised machine learning algorithms are combined with text mining techniques to produce a hybrid model which consists of Naïve Bayes and support vector machines (SVM). This is to increase the efficiency and accuracy of the results obtained and also to reduce the computational cost and complexity. The system also provides an open platform where a group of persons with a common interest can share their comments/messages and these comments classified automatically as legal or illegal. This improves the quality of conversation among users. The hybrid model was developed using WEKA tools and Java programming language. The result shows that the hybrid model gave 96.76% accuracy as against the 61.45% and 69.21% of the Naïve Bayes and SVM models respectively.
研究の動機と目的
- ソーシャルメディアプラットフォームなどのビッグデータソースからの非構造的で高速なテキストデータを分析する課題に対処すること。
- 大規模データセットにおけるテキスト分類タスクにおける計算複雑性とコストを低減すること。
- ナイーブベイズとSVMを組み合わせることで、個々の機械学習モデルを上回る分類正確性を向上させること。
- ユーザーのコメントを法的か違法かに分類するためのオープンプラットフォームを構築し、会話の質を向上させること。
提案手法
- テキスト分類のため、ナイーブベイズとサポートベクターマシン(SVM)分類器を統合してハイブリッドモデルを構築した。
- 生の非構造的テキストデータから特徴を抽出するため、テキストマイニング技術を適用して前処理を行った。
- モデルはWEKA機械学習ツールキットおよびJavaプログラミング言語を用いて実装された。
- 性能評価のため、トレーニングとテストをソーシャルメディアのテキストデータセット上で実施した。
- ハイブリッドアプローチは、両アルゴリズムの長所を活かした:ナイーブベイズは確率的分類に、SVMは高次元特徴空間に適している。
- 性能は主に正確性を指標として評価され、ハイブリッドモデルと個々のモデルを比較した。
実験結果
リサーチクエスチョン
- RQ1ナイーブベイズとSVMを組み合わせることで、単独で使用する場合と比較して、ビッグデータにおけるテキスト分類の正確性が向上するか?
- RQ2ハイブリッドモデルは、大規模テキスト分類において計算複雑性とコストをどのように低減するか?
- RQ3ハイブリッドモデルは、ソーシャルメディアプラットフォーム上のユーザー生成コンテンツを法的か違法かに効果的に分類できるか、その程度はいかほどか?
- RQ4テキストマイニング技術の統合は、特徴表現と分類性能を向上させるか?
主な発見
- ハイブリッドモデルは96.76%の分類正確性を達成し、単独のナイーブベイズ(61.45%)やSVM(69.21%)を著しく上回った。
- ハイブリッドアプローチは、両アルゴリズムの補完的長所を活かすことで、計算複雑性を低減した。
- テキストマイニング技術の統合により、特徴抽出とモデル入力の品質が向上した。
- 本システムは、ソーシャルメディア上のユーザーのコメントを自動分類するためのスケーラブルでオープンなプラットフォームを提供する。
- 結果から、ハイブリッドモデルがビッグデータの4V(ボリューム、ボリューム、バリエーション、信頼性)を効果的に処理し、高い正確性を維持できることを示した。
- 本研究は、確率的モデルとカーネルベースモデルを組み合わせることで、非構造的データにおけるテキスト分類タスクの性能が向上することを確認した。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。