Skip to main content
QUICK REVIEW

[논문 리뷰] Benchmarking the Benchmark -- Analysis of Synthetic NIDS Datasets

Siamak Layeghy, Marcus Gallagher|arXiv (Cornell University)|2021. 04. 19.
Network Security and Intrusion Detection참고 문헌 8인용 수 7
한 줄 요약

이 논문은 실제 생산 환경의 네트워크 트래픽에서 유래한 실질적인 네트워크 트래픽의 통계적 특성과 비교하여 합성 NIDS 벤치마크 데이터셋의 현실성 여부를 평가한다. 아홉 가지 트래픽 특성과 차원 축소 기법을 사용하여 합성 데이터셋과 실세계 데이터셋 간에 유의미한 분포 차이가 있음을 입증하며, 합성 데이터로 훈련된 고성능 머신러닝 모델의 일반화 능력에 의문을 제기한다.

ABSTRACT

Network Intrusion Detection Systems (NIDSs) are an increasingly important tool for the prevention and mitigation of cyber attacks. A number of labelled synthetic datasets generated have been generated and made publicly available by researchers, and they have become the benchmarks via which new ML-based NIDS classifiers are being evaluated. Recently published results show excellent classification performance with these datasets, increasingly approaching 100 percent performance across key evaluation metrics such as accuracy, F1 score, etc. Unfortunately, we have not yet seen these excellent academic research results translated into practical NIDS systems with such near-perfect performance. This motivated our research presented in this paper, where we analyse the statistical properties of the benign traffic in three of the more recent and relevant NIDS datasets, (CIC, UNSW, ...). As a comparison, we consider two datasets obtained from real-world production networks, one from a university network and one from a medium size Internet Service Provider (ISP). Our results show that the two real-world datasets are quite similar among themselves in regards to most of the considered statistical features. Equally, the three synthetic datasets are also relatively similar within their group. However, and most importantly, our results show a distinct difference of most of the considered statistical features between the three synthetic datasets and the two real-world datasets. Since ML relies on the basic assumption of training and test datasets being sampled from the same distribution, this raises the question of how well the performance results of ML-classifiers trained on the considered synthetic datasets can translate and generalise to real-world networks. We believe this is an interesting and relevant question which provides motivation for further research in this space.

연구 동기 및 목표

  • 합성 NIDS 벤치마크 데이터셋이 실세계 양성 네트워크 트래픽을 실제로 반영하고 있는지 조사하기.
  • 합성 데이터셋과 실생산 네트워크 트래픽 간의 통계적 부일치를 규명하여 모델의 일반화 능력에 제약을 줄 수 있음을 밝히기.
  • 트래픽 특성 분포 기반으로 합성 NIDS 데이터셋의 현실성 평가를 위한 방법론을 제안하기.
  • 새로운 통계적 지표를 사용하여 합성 데이터셋과 실세계 트래픽 간의 유사도를 정량화하기.
  • 데이터셋 분포 불일치로 인한 과도하게 낙관적인 성능 주장이 머신러닝 기반 NIDS 연구에서 발생할 위험을 강조하기.

제안 방법

  • 최근의 세 가지 합성 NIDS 데이터셋을 선정: UNSW-NB15, CIC-IDS2017, TON-IOT.
  • 2019년도에 수집된 대학 네트워크와 ISP의 두 개의 실세계 NetFlow/IPFIX 데이터셋을 확보.
  • 이전의 변환 작업을 활용하여 모든 데이터셋을 동일한 NetFlow/IPFIX 형식으로 변환하여 다중 데이터셋 간 비교를 가능하게 하였다.
  • 통계 분석을 위해 아홉 가지 핵심 트래픽 특성(예: 플로우 지속 시간, 패킷 수, 바이트 수)을 추출하였다.
  • 차원 축소 기법 네 가지(t-SNE, UMAP 등)를 적용하여 2차원 공간에서 특성 분포를 시각화하였다.
  • 합성 데이터셋과 실세계 데이터셋의 특성 분포 간의 분산 정도를 정량화하기 위한 통계적 지표를 제안하였다.

실험 결과

연구 질문

  • RQ1합성 NIDS 데이터셋의 양성 트래픽 통계적 특성은 실세계 생산 네트워크 트래픽과 어떻게 비교되는가?
  • RQ2플로우 수준의 특성 분포 측면에서 합성 데이터셋은 실세계 트래픽과 어느 정도 유사한가?
  • RQ3합성 데이터셋과 실세계 데이터셋 간의 차이가 다수의 통계적 특성에 걸쳐 일관된가?
  • RQ4정량적 지표가 합성 데이터셋과 실세계 양성 트래픽 간의 분포 분산 정도를 신뢰성 있게 측정할 수 있는가?
  • RQ5관찰된 분포 불일치가 합성 데이터로 훈련된 머신러닝 기반 NIDS 모델의 일반화 능력을 어느 정도 약화시키는가?

주요 결과

  • 다양한 네트워크 유형과 지리적 기원을 가진 두 실세계 데이터셋 간에 통계적 특성 분포가 뚜렷한 유사성을 보였다.
  • 세 개의 합성 데이터셋은 특성 분포 측면에서 상호 유사성을 보이며, 일관된 합성 생성 패턴을 반영하고 있음을 시사한다.
  • 분석한 아홉 가지 특성 중 대부분에서 합성 데이터셋의 양성 트래픽과 실세계 생산 트래픽 간에 유의미하고 일관된 통계적 분포 차이가 존재한다.
  • 차원 축소 시각화 결과에서 합성 데이터셋과 실세계 데이터셋이 뚜렷한 클러스터로 분리되어 분포 차이를 확인할 수 있었다.
  • 제안된 통계적 지표는 실세계 데이터셋 간의 거리가 실세계 데이터셋과 각 합성 데이터셋 간의 거리보다 훨씬 작다는 점을 정량적으로 확인하였다.
  • 결과적으로, 합성 벤치마크에서의 높은 분류 성능가 실세계 구현으로의 일반화가 불가능할 수 있으며, 이는 양성 트래픽의 분포 이탈 때문임을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.