Skip to main content
QUICK REVIEW

[论文解读] Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications

Zhanshuo Ye, Yiming Hou|arXiv (Cornell University)|Jan 20, 2026
scientometrics and bibliometrics research被引用 0
一句话总结

本研究测试大语言模型是否仅利用早期文本信息就能预测未来高影响力的统计学论文,比较多种模型并揭示提示设计与时间泛化效应。

ABSTRACT

Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in large language models (LLMs) raise the question of whether early-stage textual information can provide useful signals of long-term scientific impact. Focusing on statistical publications, we propose a flexible, text-centered framework that leverages LLMs and structured prompt design to predict highly cited papers. Specifically, we utilize information available at the time of publication, including titles, abstracts, keywords, and limited bibliographic metadata. Using a large corpus of statistical papers, we evaluate predictive performance across multiple publication periods and alternative definitions of highly cited papers. The proposed approach achieves stable and competitive performance relative to existing methods and demonstrates strong generalization over time. Textual analysis further reveals that papers predicted as highly cited concentrate on recurring topics such as causal inference and deep learning. To facilitate practical use of the proposed approach, we further develop a WeChat mini program, extit{Stat Highly Cited Papers}, which provides an accessible interface for early-stage citation impact assessment. Overall, our results provide empirical evidence that LLMs can capture meaningful early signals of long-term citation impact, while also highlighting their limitations as tools for research impact assessment.

研究动机与目标

  • 在统计学领域动机并形式化从早期文本预测高被引论文的挑战。
  • 构建一个以文本为中心、以提示驱动的框架,利用出版时的信息(题目、摘要、关键词、年份、刊物)来预测高影响力。
  • 评估跨时间的预测性能与模型在对“高被引”定义(顶百分位)的不同情况下的泛化能力。
  • 识别与预测高影响论文相关的主题信号(议题)。
  • 提供一个实用工具(微信小程序)以帮助早期阶段的引用影响评估。

提出的方法

  • 开发一个结构化的提示框架,将任务对LLMs表述为评估方法创新、长期价值和问题重要性。
  • 为每个五年窗口加入时间背景综述,反映当代趋势。
  • 使用六个参考样本(三个高被引、三个非高被引)来校准判断。
  • 将输入仅限于出版时数据(题目、摘要、关键词、年份、出版社),以避免后出版信息泄露。
  • 通过API与多种LLM交互(ChatGPT 4o mini、DeepSeek、Gemini),评估准确性、真阳性/假阳性率和效率。
  • 将任务定义为二元YES/NO预测,判断论文是否会高被引(前1%、前5%、前10%)。

实验结果

研究问题

  • RQ1LLMs 是否能够从早期文本和背景信息推断论文未来的高引用潜力?
  • RQ2在使用结构化提示预测高被引论文方面,不同LLMs有何差异?
  • RQ3预测是否随时间而保持一致,是否揭示对特定议题或方法趋势的关注?
  • RQ4在统计学领域,基于LLM的影响力预测有哪些实际限制和偏差?
  • RQ5这种方法是否能够转化为研究人员与机构可使用的工具?

主要发现

  • LLMs 在仅使用早期文本数据情况下,对预测统计学论文的高被引表现稳定且具备竞争力。
  • 预测能力随时间泛化,在不同出版时期仍具信息性。
  • 预测的高被引论文往往聚焦于因果推断与深度学习等主题。
  • 开发了一个实用的微信小程序“Stat Highly Cited Papers”,以促进早期阶段的影响力评估。
  • 研究同时揭示了基于LLM的研究影响评估的潜力与局限性。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。