Skip to main content
QUICK REVIEW

[논문 리뷰] You are your Metadata: Identification and Obfuscation of Social Media Users using Metadata Information

Beatrice Perez, Mirco Musolesi|arXiv (Cornell University)|2018. 03. 27.
Misinformation and Its Impacts인용 수 5
한 줄 요약

이 논문은 게시 시간, 기기 유형, 상호작용 패턴과 같은 메타데이터만을 사용하여 지도 학습을 통해 소셜미디어 사용자 신원을 정확하게 식별할 수 있음을 입증한다. K-Nearest Neighbors 모델은 10,000명의 사용자 중 한 명을 96.7%의 정확도로 식별하며, 상위 10명의 후보자를 고려할 경우 정확도가 99.22%로 상승한다. 또한, 메타데이터를 왜곡하는 기법들조차도 이 위험을 의미 있게 감소시키지 못한다. 즉, 데이터의 60%가 왜곡된 경우에도 마찬가지다.

ABSTRACT

Metadata are associated to most of the information we produce in our daily interactions and communication in the digital world. Yet, surprisingly, metadata are often still catergorized as non-sensitive. Indeed, in the past, researchers and practitioners have mainly focused on the problem of the identification of a user from the content of a message. In this paper, we use Twitter as a case study to quantify the uniqueness of the association between metadata and user identity and to understand the effectiveness of potential obfuscation strategies. More specifically, we analyze atomic fields in the metadata and systematically combine them in an effort to classify new tweets as belonging to an account using different machine learning algorithms of increasing complexity. We demonstrate that through the application of a supervised learning algorithm, we are able to identify any user in a group of 10,000 with approximately 96.7% accuracy. Moreover, if we broaden the scope of our search and consider the 10 most likely candidates we increase the accuracy of the model to 99.22%. We also found that data obfuscation is hard and ineffective for this type of data: even after perturbing 60% of the training data, it is still possible to classify users with an accuracy higher than 95%. These results have strong implications in terms of the design of metadata obfuscation strategies, for example for data set release, not only for Twitter, but, more generally, for most social media platforms.

연구 동기 및 목표

  • 메타데이터만으로 소셜미디어 사용자 신원을 고유하게 식별할 수 있는지 조사하기 위해.
  • 메타데이터 특징 기반으로 사용자를 분류하는 데 있어 기계학습 모델의 성능을 평가하기 위해.
  • 메타데이터 기반 사용자 재식별에 대한 데이터 왜곡 전략의 내성 여부를 평가하기 위해.
  • 오픈 데이터셋에 메타데이터를 공개할 경우 발생하는 개인정보 유출 위험을 부각하기 위해.
  • 이러한 접근 방식을 테이터와 같은 다양한 소셜미디어 플랫폼에 적용 가능한 프레임워크로 제공하기 위해.

제안 방법

  • 연구는 500만 명의 트위터 사용자 코퍼스를 활용하여 게시 시간, 기기 유형, 상호작용 빈도와 같은 원자적 메타데이터 필드를 추출하고 분석한다.
  • 다양한 기계학습 모델—다항 로지스틱 회귀, 랜덤 포레스트, K-Nearest Neighbors—를 사용하여 사용자 메타데이터 패턴 기반으로 사용자를 분류하도록 훈련한다.
  • 특징 조합을 체계적으로 평가하여 어떤 메타데이터 조합이 가장 높은 식별 정확도를 낼 수 있는지 확인한다.
  • K-Nearest Neighbors 분류기가 대규모 사용자 집단 내에서 사용자 식별에 가장 뛰어난 성능을 보였기 때문에 주요 모델로 사용된다.
  • 왜곡 기법은 훈련 데이터의 60%를 왜곡하여 익명화를 시뮬레이션하고, 식별 정확도 유지 여부를 측정함으로써 평가된다.
  • 청결한 데이터와 왜곡된 데이터 환경에서 모두 접근 방식을 평가하여 개인정보 泄露 위험을 분석한다.

실험 결과

연구 질문

  • RQ1메시지 내용에 접근하지 못한 채 메타데이터만으로 사용자 신원을 정확하게 유추할 수 있는가?
  • RQ2다양한 기계학습 모델의 성능은 메타데이터 기반 사용자 식별에 있어 어떻게 다를까?
  • RQ3데이터 왜곡이 메타데이터 기반 사용자 재식별 위험을 어느 정도 감소시키는가?
  • RQ4어떤 메타데이터 필드 또는 조합이 사용자 식별에 가장 유용한가?
  • RQ5단일 일치가 아니라 상위 N명의 후보자를 고려할 경우 식별 정확도는 어떻게 변화하는가?

주요 결과

  • K-Nearest Neighbors 모델은 메타데이터만을 사용하여 10,000명의 사용자 중 한 명을 96.7%의 정확도로 식별한다.
  • 상위 10명의 후보자들을 고려할 경우 식별 정확도는 99.22%로 상승한다.
  • 왜곡을 통해 훈련 데이터의 60%를 왜곡한 후에도 분류 정확도가 95% 이상 유지되어 왜곡 기법이 대부분 효과가 없음을 시사한다.
  • 게시 시간, 기기 유형, 상호작용 빈도 등의 메타데이터 조합은 매우 고유한 행동 서명을 생성한다.
  • 메타데이터만으로도 고정확도의 사용자 식별이 가능하므로, 메타데이터가 비민감하다는 가정은 도전받을 필요가 있다.
  • 연구 결과는 메타데이터가 주요 콘텐츠와 동일한 개인정보 보호 조치를 받아야 한다는 점을 시사한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.