Skip to main content
QUICK REVIEW

[Paper Review] MM-COVID: A Multilingual and Multimodal Data Repository for Combating COVID-19 Disinformation

Yichuan Li, Bohan Jiang|arXiv (Cornell University)|Nov 8, 2020
Misinformation and Its ImpactsSocial Sciences32 references51 citations
TL;DR

MM-COVID provides a multilingual and multi-dimensional fake news dataset for COVID-19, combining content, social engagements, and temporal data across six languages to support cross-lingual and multimodal fake news detection and mitigation.

ABSTRACT

The COVID-19 epidemic is considered as the global health crisis of the whole society and the greatest challenge mankind faced since World War Two. Unfortunately, the fake news about COVID-19 is spreading as fast as the virus itself. The incorrect health measurements, anxiety, and hate speeches will have bad consequences on people's physical health, as well as their mental health in the whole world. To help better combat the COVID-19 fake news, we propose a new fake news detection dataset MM-COVID(Multilingual and Multidimensional COVID-19 Fake News Data Repository). This dataset provides the multilingual fake news and the relevant social context. We collect 3981 pieces of fake news content and 7192 trustworthy information from English, Spanish, Portuguese, Hindi, French and Italian, 6 different languages. We present a detailed and exploratory analysis of MM-COVID from different perspectives and demonstrate the utility of MM-COVID in several potential applications of COVID-19 fake news study on multilingual and social media.

Motivation & Objective

  • Motivate the need for a multilingual, multi-dimensional COVID-19 fake news dataset to address multilinguality and social-context signals in detection.
  • Construct MM-COVID with fake/real content in six languages and rich social/contextual features.
  • Provide baseline multilingual fake news detection methods and analyze data characteristics to guide future research.

Proposed method

  • Collect veracity labels from Snopes and Poynter in English, Spanish, Portuguese, Hindi, French, and Italian.
  • Crawl source content with Newspaper3k and extract metadata (URL, language, date, text, image).
  • Gather social engagements (tweets, replies, retweets) via Twitter advanced search and twarc; collect user profiles and timelines.
  • Analyze content, language, social context, and temporal features to characterize differences between fake and real news.
  • Evaluate baseline detectors using content-only, social-context-only, and joint content+social-context models (SVM, XGBoost, dEFEND variants) across languages.

Experimental results

Research questions

  • RQ1RQ1 How do content-only, social-context-only, and joint models perform when sufficient labeled data is available across languages?
  • RQ2RQ2 How does performance change under low-resource conditions with cross-language data sharing?
  • RQ3RQ3 Can social-context signals enable cross-lingual fake news detection when a target language has no labeled data?

Key findings

  • MM-COVID enables cross-language fake news detection by combining multilingual content with social context.
  • Content+social-context models (dEFEND variants) outperform content-only baselines across languages in sufficient-resource settings.
  • In low-resource settings, social context helps when using target-language data plus auxiliary source-language data; without any target-language data, cross-lingual social-context models still offer competitive performance.
  • Temporal social-engagement patterns reveal language-invariant signals that can support early fake news detection across languages.
  • Bot-like user behavior correlates with fake news engagement in several languages, indicating the value of user-profile features for detection.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.