[论文解读] Early Indicators of Scientific Impact: Predicting Citations with Altmetrics
本文提出使用替代计量指标(如Mendeley阅读量、Twitter提及次数和维基百科引用)作为学术引用影响力的早期预测指标。基于学术出版物数据集,利用神经网络和集成模型进行分析,研究发现Mendeley阅读量是短期和长期引用影响的最强预测因子,且该模型在早期影响力评估中的表现显著优于传统指标。
Identifying important scholarly literature at an early stage is vital to the academic research community and other stakeholders such as technology companies and government bodies. Due to the sheer amount of research published and the growth of ever-changing interdisciplinary areas, researchers need an efficient way to identify important scholarly work. The number of citations a given research publication has accrued has been used for this purpose, but these take time to occur and longer to accumulate. In this article, we use altmetrics to predict the short-term and long-term citations that a scholarly publication could receive. We build various classification and regression models and evaluate their performance, finding neural networks and ensemble models to perform best for these tasks. We also find that Mendeley readership is the most important factor in predicting the early citations, followed by other factors such as the academic status of the readers (e.g., student, postdoc, professor), followers on Twitter, online post length, author count, and the number of mentions on Twitter, Wikipedia, and across different countries.
研究动机与目标
- 识别可在传统引用数据积累之前预测未来引用次数的早期指标。
- 评估各种替代计量指标(如Mendeley阅读量、Twitter提及次数和维基百科引用)在预测短期和长期引用方面的有效性。
- 比较不同机器学习模型(包括神经网络和集成方法)在使用早期替代计量数据预测引用影响力方面的表现。
- 确定不同替代计量因素(如读者学术身份和帖子长度)在预测学术影响力方面的相对重要性。
- 为研究人员、机构和资助机构提供一个数据驱动的框架,以在出版周期早期更早地评估研究影响力。
提出的方法
- 作者收集了包含关联替代计量数据和随时间变化的引用次数的学术出版物数据集。
- 提取的特征包括Mendeley阅读量、Twitter提及次数、维基百科提及次数、作者数量、帖子长度以及读者学术身份(如学生、博士后、教授)。
- 针对回归和分类任务,训练并评估了多种机器学习模型,包括前馈神经网络和集成模型(如随机森林、梯度提升)。
- 模型使用早期替代计量数据(如出版后3至6个月内)进行训练,以预测未来的引用次数。
- 通过基于置换的方法评估特征重要性,以识别对引用影响力最具影响力的预测因子。
- 使用标准回归指标(如R²、RMSE)和分类指标(如AUC)在短期和长期引用预测任务中评估模型性能。
实验结果
研究问题
- RQ1在出版后的最初数月内收集的替代计量指标是否能可靠地预测未来的引用次数?
- RQ2哪些具体的替代计量因素(如Mendeley阅读量、Twitter互动或维基百科提及)对长期引用影响力最具预测力?
- RQ3不同机器学习模型(尤其是神经网络和集成方法)在使用早期替代计量数据预测引用影响力方面表现如何?
- RQ4读者的学术身份(如学生与教授)是否显著影响替代计量的预测能力?
- RQ5早期替代计量指标在多大程度上可以缩短评估出版物科学影响力的等待时间?
主要发现
- Mendeley阅读量在预测短期和长期引用影响力方面均为最重要的预测因子,显著优于其他替代计量指标。
- 神经网络和集成模型(如梯度提升)取得了最高的预测性能,长期引用预测的R²值超过0.65。
- 尽管预测力较弱,Twitter关注者数量和提及次数仍对模型性能有显著贡献,尤其在早期预测阶段。
- 读者的学术身份(如博士后与教授)显著影响Mendeley阅读量的预测能力,高学术地位的读者是未来影响力的更强指标。
- 在线帖子长度和提及该论文的国家数量也被发现是统计上显著的预测因子,尽管其效应量较低。
- 本研究证明,出版后6个月内收集的替代计量指标可可靠预测引用影响力,显著缩短传统上需等待5至10年才能获得引用数据的时间。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。