Skip to main content
QUICK REVIEW

[논문 리뷰] CIMTDetect: A Community Infused Matrix-Tensor Coupled Factorization Based Method for Fake News Detection

Shashank Gupta, Raghuveer Thirukovalluru|arXiv (Cornell University)|2018. 09. 14.
Topic Modeling참고 문헌 28인용 수 9
한 줄 요약

CIMTDetect는 <뉴스, 사용자, 커뮤니티> 3모드 텐서를 통해 뉴스 참여 패턴을 모델링하고 에코 채터 구조를 활용하여 잠재 표현을 향상시킴으로써 가짜 뉴스 탐지에 커뮤니티를 통합한 행렬-텐서 결합 인자분해 방법을 제안한다. 이 방법은 두 개의 실세계 데이터셋에서 베이스라인을 능가하며, 뉴스 코hort 분석 및 공동 뉴스 추천과 같은 보조 작업으로의 일반화 능력도 효과적으로 보여준다.

ABSTRACT

Detecting whether a news article is fake or genuine is a crucial task in today's digital world where it's easy to create and spread a misleading news article. This is especially true of news stories shared on social media since they don't undergo any stringent journalistic checking associated with main stream media. Given the inherent human tendency to share information with their social connections at a mouse-click, fake news articles masquerading as real ones, tend to spread widely and virally. The presence of echo chambers (people sharing same beliefs) in social networks, only adds to this problem of wide-spread existence of fake news on social media. In this paper, we tackle the problem of fake news detection from social media by exploiting the very presence of echo chambers that exist within the social network of users to obtain an efficient and informative latent representation of the news article. By modeling the echo-chambers as closely-connected communities within the social network, we represent a news article as a 3-mode tensor of the structure - and propose a tensor factorization based method to encode the news article in a latent embedding space preserving the community structure. We also propose an extension of the above method, which jointly models the community and content information of the news article through a coupled matrix-tensor factorization framework. We empirically demonstrate the efficacy of our method for the task of Fake News Detection over two real-world datasets. Further, we validate the generalization of the resulting embeddings over two other auxiliary tasks, namely: extbf{1)} News Cohort Analysis and extbf{2)} Collaborative News Recommendation. Our proposed method outperforms appropriate baselines for both the tasks, establishing its generalization.

연구 동기 및 목표

  • 소셜 미디어에서의 가짜 뉴스 유포 문제를 다루며, 특히 에코 채터와 확인 편향에 의해 악화되는 문제를 해결한다.
  • 소셜 네트워크의 커뮤니티 구조를 뉴스 기사에 대한 정보적인 잠재 표현의 원천으로 활용한다.
  • 텍스트 콘텐츠와 커뮤니티 참여 패턴을 동시에 모델링하여 가짜 뉴스 탐지 성능을 향상시킨다.
  • 학습된 임베딩이 가짜 뉴스 탐지 외에도 뉴스 코hort 분석 및 공동 뉴스 추천과 같은 보조 작업으로 일반화되는지를 입증한다.
  • 더 강력하고 구분력 있는 뉴스 표현 학습을 위한 커뮤니티 구조와 콘텐츠 특징의 통합 신규 프레임워크를 구축한다.

제안 방법

  • 소셜 네트워크 커뮤니티 내의 참여 패턴을 포착하기 위해 뉴스 기사들을 <뉴스, 사용자, 커뮤니티> 차원을 가진 3모드 텐서로 표현한다.
  • 텐서 인자분해를 적용하여 커뮤니티 구조와 사용자-뉴스 상호작용을 유지하는 저차원 잠재 임베딩을 학습한다.
  • 콘텐츠 매트릭스를 통해 텍스트 콘텐츠를, 텐서를 통해 커뮤니티 참여를 모델링하는 결합 행렬-텐서 인자분해(CMTF) 프레임워크를 도입하여 접근을 확장한다.
  • 뉴스 임베딩이 언어적 특징과 커뮤니티 수준의 확산 패턴을 모두 포함하는 공유 잠재 공간을 학습한다.
  • 협업 필터링 및 클러스터링 작업에서 비교를 위해 비음수 행렬 인자분해(NMF)를 베이스라인으로 사용한다.
  • 여러 데이터셋과 구성에서 F1-score, Precision@k, 그리고 클러스터링 지표(Silhouette 및 Calinski-Harabasz 지수)를 사용하여 모델 성능을 평가한다.

실험 결과

연구 질문

  • RQ1소셜 네트워크에서 에코 채터 커뮤니티를 모델링하면 가짜 뉴스 탐지 시스템의 성능 향상에 기여하는가?
  • RQ2결합 행렬-텐서 인자분해를 통해 텍스트 콘텐츠와 커뮤니티 구조를 통합적으로 모델링할 경우, 각각의 모odal을 별도로 모델링하는 것보다 탐지 정확도가 어떻게 향상되는가?
  • RQ3학습된 잠재 임베딩이 뉴스 코hort 분석 및 공동 뉴스 추천과 같은 후행 작업으로 얼마나 잘 일반화되는가?
  • RQ4다양한 데이터셋에서 탐지 성능을 최대화하기 위해 뉴스 임베딩의 최적 차원은 어느 정도인가?
  • RQ5제안된 방법은 특히 뉴스 요인의 잠재 차원에 대해 어떤 민감도를 보이는가?

주요 결과

  • CIMTDetect는 두 개의 실세계 데이터셋에서 콘텐츠 전용 및 커뮤니티 전용 베이스라인보다 뛰어난 가짜 뉴스 탐지 성능을 달성한다.
  • F1-score 기준으로 NMF 및 기타 최첨단 방법보다도 뛰어나며, 특정 뉴스 임베딩 차원에서 가장 높은 F1-score를 기록한다.
  • CIMTDetect가 학습한 임베딩은 보조 작업으로서 잘 일반화되며, UC-NMF 베이스라인 대비 공동 뉴스 추천에서 Precision@1 및 Precision@5가 유의미하게 향상된다.
  • 클러스터링 결과는 CIMTDetect가 생성한 임베딩이 NMF 및 CITDetect보다 더 높은 Silhouette 및 Calinski-Harabasz 지수를 보이며, 더 명확하고 밀도 높은 클러스터를 생성함을 시사한다.
  • 파rameter 민감도 분석 결과, CIMTDetect와 CITDetect는 각각 다른 최적의 임베딩 차원에서 최고 성능를 기록함으로써 두 모델 간에 서로 다른 학습 역학을 가짐을 확인한다.
  • CMTF 기반 확장(CIMTDetect)은 독립적인 텐서 인자분해(CITDetect) 및 콘텐츠 전용 방법보다 일관되게 뛰어난 성능을 보이며, 콘텐츠와 커뮤니티 구조의 동시 모델링이 유의미한 이점을 제공함을 확인한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.