Skip to main content
QUICK REVIEW

[论文解读] A Survey of Information Cascade Analysis: Models, Predictions, and Recent Advances

Fan Zhou, Xovee Xu|arXiv (Cornell University)|May 22, 2020
Complex Network Analysis Techniques参考文献 191被引用 49
一句话总结

This survey comprehensively reviews information cascade popularity prediction methods, categorizing by feature-based, generative, and deep learning approaches across macro-, micro-, and meso-level predictions, and outlines open challenges.

ABSTRACT

The deluge of digital information in our daily life -- from user-generated content, such as microblogs and scientific papers, to online business, such as viral marketing and advertising -- offers unprecedented opportunities to explore and exploit the trajectories and structures of the evolution of information cascades. Abundant research efforts, both academic and industrial, have aimed to reach a better understanding of the mechanisms driving the spread of information and quantifying the outcome of information diffusion. This article presents a comprehensive review and categorization of information popularity prediction methods, from feature engineering and stochastic processes, through graph representation, to deep learning-based approaches. Specifically, we first formally define different types of information cascades and summarize the perspectives of existing studies. We then present a taxonomy that categorizes existing works into the aforementioned three main groups as well as the main subclasses in each group, and we systematically review cutting-edge research work. Finally, we summarize the pros and cons of existing research efforts and outline the open challenges and opportunities in this field.

研究动机与目标

  • Define types and problem formulations of information cascades and popularity prediction.
  • Provide a taxonomy of prediction methods across feature-based, generative, and deep learning approaches.
  • Review macro-, micro-, and meso-level prediction tasks and evaluation protocols.
  • Summarize datasets, evaluation metrics, and open challenges in information diffusion research.

提出的方法

  • Classify information cascades into prediction as classification or regression.
  • Differentiate ex-ante prediction from peeking strategies based on available data.
  • Organize methods into feature-based, generative, and deep learning categories with cross-network applicability.
  • Discuss evaluation metrics and benchmark datasets used in popularity prediction.
  • Highlight the trade-offs, advantages, and limitations of each methodological approach.
  • Survey recent literature on graph representation learning and sequential models for diffusion tasks.

实验结果

研究问题

  • RQ1What are the main problem formulations for information cascade popularity prediction (classification vs regression, ex-ante vs peeking, macro/micro/meso levels)?
  • RQ2What are the predominant methodological approaches (feature-based, generative, deep learning) and their trade-offs across different networks and data types?
  • RQ3How do evaluation metrics and datasets shape comparisons of information cascade models?
  • RQ4What are the open challenges and opportunities in modeling information diffusion and predicting cascade popularity?

主要发现

  • The paper provides a broad taxonomy: prediction can be classification or regression, before or after publication, and macro-, micro-, or meso-level in scope.
  • It covers three method groups: feature-based methods, generative models, and deep learning approaches, including graph representation learning and sequential models.
  • It discusses multiple networks and data domains (e.g., social networks, content sharing, and citation networks) and notes deep learning methods have recently become more popular.
  • It reviews evaluation metrics (e.g., accuracy, precision, recall, F1, AUC, MAE, RMSE) and discusses issues with highly skewed popularity distributions.
  • It emphasizes that many models are not easily generalizable across platforms and datasets, and highlights open challenges and opportunities in the field.

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。