[Paper Review] Early Indicators of Scientific Impact: Predicting Citations with Altmetrics
This paper proposes using altmetrics—such as Mendeley readership, Twitter mentions, and Wikipedia references—as early predictors of scholarly citation impact. Using neural networks and ensemble models on a dataset of academic publications, the authors find that Mendeley readership is the strongest predictor of both short- and long-term citations, with model performance significantly outperforming traditional metrics in early impact assessment.
Identifying important scholarly literature at an early stage is vital to the academic research community and other stakeholders such as technology companies and government bodies. Due to the sheer amount of research published and the growth of ever-changing interdisciplinary areas, researchers need an efficient way to identify important scholarly work. The number of citations a given research publication has accrued has been used for this purpose, but these take time to occur and longer to accumulate. In this article, we use altmetrics to predict the short-term and long-term citations that a scholarly publication could receive. We build various classification and regression models and evaluate their performance, finding neural networks and ensemble models to perform best for these tasks. We also find that Mendeley readership is the most important factor in predicting the early citations, followed by other factors such as the academic status of the readers (e.g., student, postdoc, professor), followers on Twitter, online post length, author count, and the number of mentions on Twitter, Wikipedia, and across different countries.
Motivation & Objective
- To identify early indicators of scientific impact that can predict future citation counts before traditional citation data accumulates.
- To evaluate the effectiveness of various altmetrics—such as Mendeley readership, Twitter mentions, and Wikipedia references—in forecasting short- and long-term citations.
- To compare the performance of different machine learning models, including neural networks and ensemble methods, in predicting citation impact using early altmetric data.
- To determine the relative importance of different altmetric factors, such as reader academic status and post length, in predicting scholarly impact.
- To provide a data-driven framework for researchers, institutions, and funding bodies to assess research impact earlier in the publication lifecycle.
Proposed method
- The authors collect a dataset of scholarly publications with associated altmetric data and citation counts over time.
- They extract features including Mendeley readership, Twitter mentions, Wikipedia mentions, author count, post length, and reader academic status (e.g., student, postdoc, professor).
- They train and evaluate multiple machine learning models, including feedforward neural networks and ensemble models (e.g., random forests, gradient boosting), for both regression and classification tasks.
- The models are trained on early altmetric data (e.g., within the first 3–6 months post-publication) to predict future citation counts.
- Feature importance is assessed using permutation-based methods to identify the most influential predictors of citation impact.
- Performance is evaluated using standard regression metrics (e.g., R², RMSE) and classification metrics (e.g., AUC) across short- and long-term citation prediction tasks.
Experimental results
Research questions
- RQ1Can altmetrics collected within the first few months of publication reliably predict future citation counts?
- RQ2Which specific altmetric factors—such as Mendeley readership, Twitter engagement, or Wikipedia mentions—are most predictive of long-term citation impact?
- RQ3How do different machine learning models, particularly neural networks and ensemble methods, compare in predicting citation impact using early altmetric data?
- RQ4Does the academic status of readers (e.g., student vs. professor) significantly influence the predictive power of altmetrics?
- RQ5To what extent can early altmetric indicators reduce the time needed to assess the scientific impact of a publication?
Key findings
- Mendeley readership emerged as the most important predictor of both short- and long-term citation impact, significantly outperforming other altmetric indicators.
- Neural networks and ensemble models (e.g., gradient boosting) achieved the highest predictive performance, with R² values exceeding 0.65 for long-term citation prediction.
- Twitter followers and mentions, while less predictive, still contributed meaningfully to model performance, especially in early-stage prediction.
- The academic status of readers (e.g., postdoc vs. professor) significantly influenced the predictive power of Mendeley readership, with higher-status readers being stronger indicators of future impact.
- Online post length and the number of countries mentioning a paper were also found to be statistically significant predictors, though with lower effect sizes.
- The study demonstrates that altmetrics collected within the first 6 months post-publication can reliably forecast citation impact, reducing the traditional 5–10 year wait for citation data.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.