Skip to main content
QUICK REVIEW

[Paper Review] A Survey of Information Cascade Analysis: Models, Predictions, and Recent Advances

Fan Zhou, Xovee Xu|arXiv (Cornell University)|May 22, 2020
Complex Network Analysis TechniquesPhysics and Astronomy191 references49 citations
TL;DR

This survey comprehensively reviews information cascade popularity prediction methods, categorizing by feature-based, generative, and deep learning approaches across macro-, micro-, and meso-level predictions, and outlines open challenges.

ABSTRACT

The deluge of digital information in our daily life -- from user-generated content, such as microblogs and scientific papers, to online business, such as viral marketing and advertising -- offers unprecedented opportunities to explore and exploit the trajectories and structures of the evolution of information cascades. Abundant research efforts, both academic and industrial, have aimed to reach a better understanding of the mechanisms driving the spread of information and quantifying the outcome of information diffusion. This article presents a comprehensive review and categorization of information popularity prediction methods, from feature engineering and stochastic processes, through graph representation, to deep learning-based approaches. Specifically, we first formally define different types of information cascades and summarize the perspectives of existing studies. We then present a taxonomy that categorizes existing works into the aforementioned three main groups as well as the main subclasses in each group, and we systematically review cutting-edge research work. Finally, we summarize the pros and cons of existing research efforts and outline the open challenges and opportunities in this field.

Motivation & Objective

  • Define types and problem formulations of information cascades and popularity prediction.
  • Provide a taxonomy of prediction methods across feature-based, generative, and deep learning approaches.
  • Review macro-, micro-, and meso-level prediction tasks and evaluation protocols.
  • Summarize datasets, evaluation metrics, and open challenges in information diffusion research.

Proposed method

  • Classify information cascades into prediction as classification or regression.
  • Differentiate ex-ante prediction from peeking strategies based on available data.
  • Organize methods into feature-based, generative, and deep learning categories with cross-network applicability.
  • Discuss evaluation metrics and benchmark datasets used in popularity prediction.
  • Highlight the trade-offs, advantages, and limitations of each methodological approach.
  • Survey recent literature on graph representation learning and sequential models for diffusion tasks.

Experimental results

Research questions

  • RQ1What are the main problem formulations for information cascade popularity prediction (classification vs regression, ex-ante vs peeking, macro/micro/meso levels)?
  • RQ2What are the predominant methodological approaches (feature-based, generative, deep learning) and their trade-offs across different networks and data types?
  • RQ3How do evaluation metrics and datasets shape comparisons of information cascade models?
  • RQ4What are the open challenges and opportunities in modeling information diffusion and predicting cascade popularity?

Key findings

  • The paper provides a broad taxonomy: prediction can be classification or regression, before or after publication, and macro-, micro-, or meso-level in scope.
  • It covers three method groups: feature-based methods, generative models, and deep learning approaches, including graph representation learning and sequential models.
  • It discusses multiple networks and data domains (e.g., social networks, content sharing, and citation networks) and notes deep learning methods have recently become more popular.
  • It reviews evaluation metrics (e.g., accuracy, precision, recall, F1, AUC, MAE, RMSE) and discusses issues with highly skewed popularity distributions.
  • It emphasizes that many models are not easily generalizable across platforms and datasets, and highlights open challenges and opportunities in the field.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.