Skip to main content
QUICK REVIEW

[논문 리뷰] Using BERT Encoding to Tackle the Mad-lib Attack in SMS Spam Detection.

Sergio Rojas Galeano|arXiv (Cornell University)|2021. 07. 13.
Spam and Phishing Detection참고 문헌 23인용 수 10
한 줄 요약

이 논문은 스팸 메시지가 동의어 치환을 통해 은폐되는 '마드리브 공격'에 대해 BERT의 내성적 저항성을 조사한다. 5,572건의 SMS 스팸 메시지로 구성된 데이터셋을 사용하여, 메시지당 평균 1.82개의 단어가 교체된 상황에서도 BERT는 여전히 96%의 균형 정확도를 유지하는 것으로 나타났다. 반면 기존 모델(BoW 및 TFIDF)은 우연수준의 성능으로 악화되며, BERT의 동의어 기반 회피 공격에 대한 우월한 의미적 내성성을 입증한다.

ABSTRACT

One of the stratagems used to deceive spam filters is to substitute vocables with synonyms or similar words that turn the message unrecognisable by the detection algorithms. In this paper we investigate whether the recent development of language models sensitive to the semantics and context of words, such as Google's BERT, may be useful to overcome this adversarial attack (called as per the word substitution game). Using a dataset of 5572 SMS spam messages, we first established a baseline of detection performance using widely known document representation models (BoW and TFIDF) and the novel BERT model, coupled with a variety of classification algorithms (Decision Tree, kNN, SVM, Logistic Regression, Naive Bayes, Multilayer Perceptron). Then, we built a thesaurus of the vocabulary contained in these messages, and set up a Mad-lib attack experiment in which we modified each message of a held out subset of data (not used in the baseline experiment) with different rates of substitution of original words with synonyms from the thesaurus. Lastly, we evaluated the detection performance of the three representation models (BoW, TFIDF and BERT) coupled with the best classifier from the baseline experiment (SVM). We found that the classic models achieved a 94% Balanced Accuracy (BA) in the original dataset, whereas the BERT model obtained 96%. On the other hand, the Mad-lib attack experiment showed that BERT encodings manage to maintain a similar BA performance of 96% with an average substitution rate of 1.82 words per message, and 95% with 3.34 words substituted per message. In contrast, the BA performance of the BoW and TFIDF encoders dropped to chance. These results hint at the potential advantage of BERT models to combat these type of ingenious attacks, offsetting to some extent for the inappropriate use of semantic relationships in language.

연구 동기 및 목표

  • BERT 임베딩이 악성 동의어 치환 공격 하에서 SMS 스팸 탐지에 얼마나 효과적인지 평가하는 것.
  • 이러한 공격 조건에서 BERT의 성능을 기존 텍스트 표현 모델(BoW 및 TFIDF)과 비교하는 것.
  • BERT와 같은 문맥 기반 언어 모델이 동의어 치환을 통해 의미가 변형된 메시지에서도 높은 탐지 정확도를 유지할 수 있는지 평가하는 것.
  • BERT가 스팸 탐지에서 악성 변형을 처리할 수 있는 내성적 한계를 규명하는 것.

제안 방법

  • SMS 스팸 데이터셋의 어휘에서 사전을 구축하여 제어된 동의어 치환을 가능하게 하였다.
  • 기본 성능를 확보하기 위해 BoW, TFIDF, BERT 임베딩에 대해 다수의 분류기(SVM, 로지스틱 회귀 등)를 훈련 및 평가하였다.
  • 마드리브 공격을 시뮬레이션하기 위해 각 메시지에 다양한 비율의 동의어 치환을 가한 검증용 테스트 세트를 생성하였다.
  • 모든 모델 및 공격 수준에서 균형 정확도(BA)를 측정하여 탐지 성능을 측정하였다.
  • 공격 평가 단계에서 사용하기 위해 기본 실험에서 가장 높은 성능를 보인 SVM을 선정하였다.
  • 마드리브 공격에서 점차 증가하는 치환 비율에 따라 각 표현 모델(BoW, TFIDF, BERT)의 강건성을 평가하였다.

실험 결과

연구 질문

  • RQ1메시지가 동의어 치환을 통해 변형될 경우(BERT 임베딩이 마드리브 공격에 의해 영향을 받는가?
  • RQ2BoW 및 TFIDF와 같은 기존 모델의 성능은 BERT에 비해 동의어 기반 악성 공격에서 어떻게 악화되는가?
  • RQ3BERT의 탐지 성능가 중대하게 저하되기 시작하는 평균 치환 비율은 어느 수준인가?
  • RQ4BERT가 여전히 고전적 모델보다 스팸 탐지에서 뛰어난 성능를 유지할 수 있는 의미적 변형의 한계는 어디인가?

주요 결과

  • 원본 데이터셋에서 BERT는 96%의 균형 정확도를 기록하여 BoW 및 TFIDF(94%)를 능가하였다.
  • 메시지당 평균 1.82개의 단어가 교체된 마드리브 공격 조건에서도 BERT는 여전히 96%의 균형 정확도를 유지하였다.
  • 메시지당 평균 3.34개의 단어가 교체된 상황에서도 BERT의 균형 정확도는 95%를 유지하여 강력한 내성성을 보였다.
  • 반면, BoW 및 TFIDF 모델의 균형 정확도는 동일한 공격 조건에서 우연수준으로 떨어졌다.
  • 결과는 BERT의 문맥 이해 능력이 동의어 치환을 통한 의미적 은폐에도 불구하고 스팸을 인식할 수 있음을 시사한다.
  • 본 연구는 BERT 모델이 기존의 백오브워드 및 TFIDF 표현 방식보다 동의어 기반 악성 공격에 더 강건함을 입증한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.