[论文解读] A Survey of Information Cascade Analysis: Models, Predictions, and Recent Advances
This survey comprehensively reviews information cascade popularity prediction methods, categorizing by feature-based, generative, and deep learning approaches across macro-, micro-, and meso-level predictions, and outlines open challenges.
The deluge of digital information in our daily life -- from user-generated content, such as microblogs and scientific papers, to online business, such as viral marketing and advertising -- offers unprecedented opportunities to explore and exploit the trajectories and structures of the evolution of information cascades. Abundant research efforts, both academic and industrial, have aimed to reach a better understanding of the mechanisms driving the spread of information and quantifying the outcome of information diffusion. This article presents a comprehensive review and categorization of information popularity prediction methods, from feature engineering and stochastic processes, through graph representation, to deep learning-based approaches. Specifically, we first formally define different types of information cascades and summarize the perspectives of existing studies. We then present a taxonomy that categorizes existing works into the aforementioned three main groups as well as the main subclasses in each group, and we systematically review cutting-edge research work. Finally, we summarize the pros and cons of existing research efforts and outline the open challenges and opportunities in this field.
研究动机与目标
- Define types and problem formulations of information cascades and popularity prediction.
- Provide a taxonomy of prediction methods across feature-based, generative, and deep learning approaches.
- Review macro-, micro-, and meso-level prediction tasks and evaluation protocols.
- Summarize datasets, evaluation metrics, and open challenges in information diffusion research.
提出的方法
- Classify information cascades into prediction as classification or regression.
- Differentiate ex-ante prediction from peeking strategies based on available data.
- Organize methods into feature-based, generative, and deep learning categories with cross-network applicability.
- Discuss evaluation metrics and benchmark datasets used in popularity prediction.
- Highlight the trade-offs, advantages, and limitations of each methodological approach.
- Survey recent literature on graph representation learning and sequential models for diffusion tasks.
实验结果
研究问题
- RQ1What are the main problem formulations for information cascade popularity prediction (classification vs regression, ex-ante vs peeking, macro/micro/meso levels)?
- RQ2What are the predominant methodological approaches (feature-based, generative, deep learning) and their trade-offs across different networks and data types?
- RQ3How do evaluation metrics and datasets shape comparisons of information cascade models?
- RQ4What are the open challenges and opportunities in modeling information diffusion and predicting cascade popularity?
主要发现
- The paper provides a broad taxonomy: prediction can be classification or regression, before or after publication, and macro-, micro-, or meso-level in scope.
- It covers three method groups: feature-based methods, generative models, and deep learning approaches, including graph representation learning and sequential models.
- It discusses multiple networks and data domains (e.g., social networks, content sharing, and citation networks) and notes deep learning methods have recently become more popular.
- It reviews evaluation metrics (e.g., accuracy, precision, recall, F1, AUC, MAE, RMSE) and discusses issues with highly skewed popularity distributions.
- It emphasizes that many models are not easily generalizable across platforms and datasets, and highlights open challenges and opportunities in the field.
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。