Skip to main content
QUICK REVIEW

[Paper Review] Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications

Zhanshuo Ye, Yiming Hou|arXiv (Cornell University)|Jan 20, 2026
scientometrics and bibliometrics research0 citations
TL;DR

The study tests whether LLMs can predict future high-impact statistical papers using only early textual information, comparing multiple models and revealing prompts-design and temporal generalization effects.

ABSTRACT

Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in large language models (LLMs) raise the question of whether early-stage textual information can provide useful signals of long-term scientific impact. Focusing on statistical publications, we propose a flexible, text-centered framework that leverages LLMs and structured prompt design to predict highly cited papers. Specifically, we utilize information available at the time of publication, including titles, abstracts, keywords, and limited bibliographic metadata. Using a large corpus of statistical papers, we evaluate predictive performance across multiple publication periods and alternative definitions of highly cited papers. The proposed approach achieves stable and competitive performance relative to existing methods and demonstrates strong generalization over time. Textual analysis further reveals that papers predicted as highly cited concentrate on recurring topics such as causal inference and deep learning. To facilitate practical use of the proposed approach, we further develop a WeChat mini program, extit{Stat Highly Cited Papers}, which provides an accessible interface for early-stage citation impact assessment. Overall, our results provide empirical evidence that LLMs can capture meaningful early signals of long-term citation impact, while also highlighting their limitations as tools for research impact assessment.

Motivation & Objective

  • Motivate and formalize the challenge of predicting highly cited papers from early-stage text in statistics.
  • Build a text-centered, prompt-driven framework that uses publication-time information (title, abstract, keywords, year, venue) to predict high impact.
  • Evaluate cross-time predictive performance and model generalization across different definitions of “highly cited” (top percentiles).
  • Identify thematic signals (topics) associated with predicted high-impact papers.
  • Provide a practical tool (WeChat mini program) to aid early-stage citation impact assessment.

Proposed method

  • Develop a structured prompting framework that frames the task for LLMs as evaluating methodological innovation, long-term value, and problem significance.
  • Incorporate a temporal background summary for each five-year window to reflect contemporaneous trends.
  • Use six reference examples (three highly cited, three not) to calibrate judgments.
  • Restrict inputs to publication-time data (title, abstract, keywords, year, publisher) to avoid leakage of post-publication information.
  • Interact with multiple LLMs via APIs (ChatGPT 4o mini, DeepSeek, Gemini) and assess accuracy, true/false positive rates, and efficiency.
  • Define the task as a binary YES/NO prediction for whether a paper will be highly cited (top 1%, 5%, or 10%).

Experimental results

Research questions

  • RQ1Can LLMs infer a paper’s future high citation potential from early textual and contextual information?
  • RQ2How do different LLMs compare in predicting highly cited papers using structured prompts?
  • RQ3Are predictions consistent over time, and do they reveal attention to particular topics or methodological trends?
  • RQ4What are the practical limits and biases of LLM-based impact prediction in statistics?
  • RQ5Can the approach be translated into usable tools for researchers and institutions?

Key findings

  • LLMs show stable and competitive performance in predicting highly cited statistical papers using only early textual data.
  • Prediction ability generalizes over time, remaining informative across publication periods.
  • Predicted high-impact papers tend to cluster around topics like causal inference and deep learning.
  • A practical WeChat mini program (Stat Highly Cited Papers) was developed to facilitate early-stage impact assessment.
  • The study highlights both the potential and the limitations of LLM-based research impact assessment.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.