Skip to main content
QUICK REVIEW

[論文レビュー] Improving the methods of email classification based on words ontology

Foruzan Kiamarzpour, Rouhollah Dianat|arXiv (Cornell University)|Oct 22, 2013
Spam and Phishing Detection参考文献 10被引用数 3
ひとこと要約

本論文は、アンサンブル意思決定木と語彙的オントロジーを統合することで、精度を向上させる新しいメールスパム分類手法を提案する。語彙的オントロジーからの意味的関係を活用して特徴選択を精緻化し、投票による複数の意思決定木を組み合わせることで、従来のキーワードベース手法に比べてスパム検出性能が向上する。

ABSTRACT

The Internet has dramatically changed the relationship among people and their relationships with others people and made the valuable information available for the users. Email is the service, which the Internet provides today for its own users; this service has attracted most of the users' attention due to the low cost. Along with the numerous benefits of Email, one of the weaknesses of this service is that the number of received emails is continually being enhanced, thus the ways are needed to automatically filter these disturbing letters. Most of these filters utilize a combination of several techniques such as the Black or white List, using the keywords and so on in order to identify the spam more accurately In this paper, we introduce a new method to classify the spam. We are seeking to increase the accuracy of Email classification by combining the output of several decision trees and the concept of ontology.

研究の動機と目的

  • 増加するスパムの量と複雑さに起因するメールスパムの深刻な課題に対処すること。
  • 従来のキーワードベースのフィルタリング手法を上回る分類精度を向上させること。
  • 語彙的オントロジーからの意味的知識を機械学習モデルに統合し、より良い特徴表現を実現すること。
  • 意思決定木とオントロジー推論を組み合わせたハイブリッド手法を開発し、頑健なスパム検出を実現すること。

提案手法

  • メール用語の意味的関係を捉える分野特化型語彙的オントロジーを構築する。
  • オントロジーを活用して特徴抽出を強化し、生のキーワードを意味的に意味のある属性に変換する。
  • オントロジー強化特徴集合上で複数の意思決定木を学習させ、一般化性能を向上させる。
  • 個々の意思決定木の出力を投票機構を用いて統合し、最終的な分類を出力する。
  • オントロジーからの意味的類似度測定を適用して曖昧さを解消し、誤検出を低減する。
  • 交差検証を用いてアンサンブルモデルを最適化し、適合度と再現率のバランスを取る。

実験結果

リサーチクエスチョン

  • RQ1語彙的オントロジーをメール分類に統合することで、従来のキーワード手法に比べて検出精度が向上するか?
  • RQ2特徴の意味的強化は、スパムフィルタリングにおける意思決定木モデルの性能にどのように影響するか?
  • RQ3複数の意思決定木を組み合わせることで、スパム検出における頑健性がどの程度向上し、過学習がどの程度低減されるか?
  • RQ4オントロジーに基づく特徴選択は、メール分類における誤検出の低減にどのような影響を及えるか?

主な発見

  • 提案手法は、ベースラインのキーワードベース手法に比べて分類精度が顕著に向上した。
  • 語彙的オントロジーの統合により、メールコンテンツの意味的理解が向上し、誤検出率が低下した。
  • 意思決定木のアンサンブルは単一木モデルを上回り、未学習データに対する一般化性能が優れていた。
  • 意味的特徴強化により、多様なスパムカテゴリにわたり一貫した性能が得られた。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。