[논문 리뷰] From Threat Reports to Continuous Threat Intelligence: A Comparison of Attack Technique Extraction Methods from Textual Artifacts
이 연구는 TF-IDF, LSI 및 기타 NLP 기법을 사용하여 사이버보안 위협 보고서에서 기반 텍스트 TTP(전술, 기법, 절차) 추출 방법 다섯 가지를 평가하고 비교한다. TF-IDF와 LSI가 각각 84% 및 83%의 F1 스코어로 가장 높은 성능을 기록했으며, 불균형 데이터셋에서 오버샘플링 기법이 성능 향상에 기여함을 입증하여 향후 TTP 추출 연구의 기준이 되는 데이터셋과 오픈소스 코드를 제공한다.
The cyberthreat landscape is continuously evolving. Hence, continuous monitoring and sharing of threat intelligence have become a priority for organizations. Threat reports, published by cybersecurity vendors, contain detailed descriptions of attack Tactics, Techniques, and Procedures (TTP) written in an unstructured text format. Extracting TTP from these reports aids cybersecurity practitioners and researchers learn and adapt to evolving attacks and in planning threat mitigation. Researchers have proposed TTP extraction methods in the literature, however, not all of these proposed methods are compared to one another or to a baseline. extit{The goal of this study is to aid cybersecurity researchers and practitioners choose attack technique extraction methods for monitoring and sharing threat intelligence by comparing the underlying methods from the TTP extraction studies in the literature.} In this work, we identify ten existing TTP extraction studies from the literature and implement five methods from the ten studies. We find two methods, based on Term Frequency-Inverse Document Frequency(TFIDF) and Latent Semantic Indexing (LSI), outperform the other three methods with a F1 score of 84\% and 83\%, respectively. We observe the performance of all methods in F1 score drops in the case of increasing the class labels exponentially. We also implement and evaluate an oversampling strategy to mitigate class imbalance issues. Furthermore, oversampling improves the classification performance of TTP extraction. We provide recommendations from our findings for future cybersecurity researchers, such as the construction of a benchmark dataset from a large corpus; and the selection of textual features of TTP. Our work, along with the dataset and implementation source code, can work as a baseline for cybersecurity researchers to test and compare the performance of future TTP extraction methods.
연구 동기 및 목표
- 사이버보안 위협 보고서에서 기존 TTP 추출 방법의 성능을 비교하여 연구자 및 실무자가 방법 선택에 참고할 수 있도록 유도하기.
- 클래스 불균형과 레이블 수 증가가 TTP 분류 성능에 미치는 영향 평가하기.
- 불균형 데이터셋으로 인한 성능 저하를 완화하기 위한 오버샘플링 기법의 효과성 평가하기.
- 다섯 가지 대표적 방법을 구현하고 평가하여 향후 TTP 추출 연구를 위한 재현 가능한 기준 마련하기.
- TTP 추출 작업에서의 특징 선택 및 모델 튜닝에 대한 최선의 실천 방안 제안하기.
제안 방법
- 기존 문헌에서 제안한 다섯 가지 TTP 추출 방법(M:LSI, M:TFIDF-NP, M:TFIDF, M:LIS-Co, M:BM25)을 구현하였으며, 지표로는 ATT&CK 프레임워크를 사용하였다.
- 위협 보고서의 공격 절차 기술 문장을 TF-IDF, LSI 및 기타 NLP 기반 특징 추출 기법을 사용해 수치적 형태로 변환하였다.
- 다섯 가지 분류기(6개)를 사용하여 기본 하이퍼파라미터 설정으로 다양한 방법 간의 분류 성능 평가를 수행하였다.
- 데이터셋의 클래스 불균형 문제를 해결하기 위해 수치적 텍스트 특징에 대해 합성 오버샘플링 전략(SMOTE)을 적용하였다.
- 변화하는 위협 환경을 시뮬레이션하기 위해 분류 레이블 수(즉, TTP 클래스 수)를 점차 증가시키며 성능을 평가하였다.
- 위협 보고서에서 기반 데이터셋을 구축하였으며, 모든 구현 및 소스 코드는 GitHub에 공개하였다.
실험 결과
연구 질문
- RQ1RQ1: 다양한 분류기에서 공격 절차의 텍스트 기술을 공격 기법으로 분류하는 데 있어 TTP 추출 방법의 성능는 어떻게 되는가?
- RQ2RQ2: 클래스 불균형은 TTP 추출 방법의 성능에 어떤 영향을 미치며, 오버샘플링은 분류 결과를 향상시키는가?
- RQ3RQ3: 분류 레이블 수가 증가함에 따라 TTP 추출 방법의 성능는 어떻게 저하되는가?
- RQ4RQ4: 어떤 텍스트 특징이 TTP 분류에 가장 효과적인가? 향후 특징 공학에 대한 함의는 무엇인가?
주요 결과
- TF-IDF 및 LSI 기반 방법이 각각 84% 및 83%의 F1 스코어로 가장 높은 성능을 기록하여 다른 방법들을 압도하였다.
- 모든 방법에서 분류 레이블 수가 증가함에 따라 성능이 저하되었으며, 이는 점차 확장되는 TTP 분류 체계에서의 확장성 문제를 시사한다.
- SMOTE를 사용한 오버샘플링 기법이 불균형 데이터셋에서 다수 클래스에 대한 편향을 줄여 분류 성능을 향상시켰다.
- 연구에서는 TF-IDF가 TTP 분류에 가장 우세한 특징로 규명되었으며, 향후 특징 공학에서 우선적으로 고려되어야 함을 시사한다.
- 저자들은 표준화된 기준 데이터셋을 구축하고 하이퍼파라미터 튜닝을 수행하여 모델 성능을 추가로 향상시킬 것을 권고한다.
- 모든 다섯 가지 방법의 데이터셋 및 소스 코드는 공개되어 있어 재현성과 향후 방법 비교를 가능하게 한다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.