Skip to main content
QUICK REVIEW

[論文レビュー] Benchmarking the Benchmark -- Analysis of Synthetic NIDS Datasets

Siamak Layeghy, Marcus Gallagher|arXiv (Cornell University)|Apr 19, 2021
Network Security and Intrusion Detection参考文献 8被引用数 7
ひとこと要約

本稿は、生産環境からの実世界のネットワークトラフィックと比較することで、合成NIDSベンチマークデータセットの現実性を評価する。9つのトラフィック特徴量と次元削減を用いて、合成データと実世界データの間で顕著な分布差が存在することを示し、合成データで訓練された高精度な機械学習モデルの一般化可能性に疑問を呈する。

ABSTRACT

Network Intrusion Detection Systems (NIDSs) are an increasingly important tool for the prevention and mitigation of cyber attacks. A number of labelled synthetic datasets generated have been generated and made publicly available by researchers, and they have become the benchmarks via which new ML-based NIDS classifiers are being evaluated. Recently published results show excellent classification performance with these datasets, increasingly approaching 100 percent performance across key evaluation metrics such as accuracy, F1 score, etc. Unfortunately, we have not yet seen these excellent academic research results translated into practical NIDS systems with such near-perfect performance. This motivated our research presented in this paper, where we analyse the statistical properties of the benign traffic in three of the more recent and relevant NIDS datasets, (CIC, UNSW, ...). As a comparison, we consider two datasets obtained from real-world production networks, one from a university network and one from a medium size Internet Service Provider (ISP). Our results show that the two real-world datasets are quite similar among themselves in regards to most of the considered statistical features. Equally, the three synthetic datasets are also relatively similar within their group. However, and most importantly, our results show a distinct difference of most of the considered statistical features between the three synthetic datasets and the two real-world datasets. Since ML relies on the basic assumption of training and test datasets being sampled from the same distribution, this raises the question of how well the performance results of ML-classifiers trained on the considered synthetic datasets can translate and generalise to real-world networks. We believe this is an interesting and relevant question which provides motivation for further research in this space.

研究の動機と目的

  • 合成NIDSベンチマークデータセットが実世界の健全なネットワークトラフィックを現実的に再現しているかどうかを調査すること。
  • 合成データセットと実生産ネットワークトラフィックの間の統計的乖離を同定し、モデルの一般化を制限する要因を特定すること。
  • トラフィック特徴量の分布に基づいて、合成NIDSデータセットの現実性を評価するための手法を提案すること。
  • 新しい統計的指標を用いて、合成データセットと実世界トラフィックの類似度を定量化すること。
  • データセットの分布不一致に起因する、機械学習ベースのNIDS研究における楽観的すぎるパフォーマンス主張のリスクを強調すること。

提案手法

  • 最近の3つの合成NIDSデータセット(UNSW-NB15、CIC-IDS2017、TON-IOT)を選定した。
  • 2019年における大学ネットワークおよびISPのNetFlow/IPFIXデータセットを2つ収集した。
  • 過去の変換作業を活用して、すべてのデータセットを共通のNetFlow/IPFIX形式に変換し、データセット間の比較を可能にした。
  • 統計的分析のため、9つの重要なトラフィック特徴量(例:フロー持続時間、パケット数、バイト数)を抽出した。
  • 4つの次元削減技術(例:t-SNE、UMAP)を適用し、2次元空間における特徴量分布の可視化を行った。
  • 合成データと実世界データの特徴量分布の乖離を定量化するための統計的指標を提案した。

実験結果

リサーチクエスチョン

  • RQ1合成NIDSデータセットにおける健全なトラフィックの統計的特徴量は、実世界の生産ネットワークトラフィックとどのように比較されるか?
  • RQ2フローレベルの特徴量分布の観点から、合成データセットはどの程度実世界トラフィックに類似しているか?
  • RQ3合成データセットと実世界データセットの間の差異は、複数の統計的特徴量にわたって一貫しているか?
  • RQ4定量的指標は、合成データと実世界の健全なトラフィックの分布乖離を信頼性高く捉えることができるか?
  • RQ5観察された分布不一致は、合成データで訓練された機械学習ベースのNIDSモデルの一般化可能性をどの程度損なっているか?

主な発見

  • 異なるネットワークタイプと地理的起源を持つにもかかわらず、2つの実世界データセットは統計的特徴量分布において強く類似している。
  • 3つの合成データセットは、特徴量分布において内部的に類似しており、一貫した合成生成パターンを示している。
  • 分析対象の9つの特徴量のほとんどにおいて、合成データの健全なトラフィックと実世界の生産トラフィックとの間に顕著で一貫した統計的乖離が存在する。
  • 次元削減の可視化により、合成データと実世界データが明確に分離されたクラスタを形成しており、分布の違いを確認できる。
  • 提案された統計的指標は、定量的に実世界データセット間の距離が、いずれの合成データセットとの距離よりも顕著に小さいことを確認した。
  • 結果から、合成ベンチマークで高い分類性能を示すことは、実世界の展開において一般化されない可能性があることが示唆される。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。