Skip to main content
QUICK REVIEW

[논문 리뷰] Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications

Zhanshuo Ye, Yiming Hou|arXiv (Cornell University)|2026. 01. 20.
scientometrics and bibliometrics research인용 수 0
한 줄 요약

본 연구는 초기 텍스트 정보만으로 LLM이 미래의 영향력 높은 통계 논문을 예측할 수 있는지 테스트하고, 여러 모델을 비교하며 프롬프트 설계 및 시간 일반화 효과를 밝힌다.

ABSTRACT

Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in large language models (LLMs) raise the question of whether early-stage textual information can provide useful signals of long-term scientific impact. Focusing on statistical publications, we propose a flexible, text-centered framework that leverages LLMs and structured prompt design to predict highly cited papers. Specifically, we utilize information available at the time of publication, including titles, abstracts, keywords, and limited bibliographic metadata. Using a large corpus of statistical papers, we evaluate predictive performance across multiple publication periods and alternative definitions of highly cited papers. The proposed approach achieves stable and competitive performance relative to existing methods and demonstrates strong generalization over time. Textual analysis further reveals that papers predicted as highly cited concentrate on recurring topics such as causal inference and deep learning. To facilitate practical use of the proposed approach, we further develop a WeChat mini program, extit{Stat Highly Cited Papers}, which provides an accessible interface for early-stage citation impact assessment. Overall, our results provide empirical evidence that LLMs can capture meaningful early signals of long-term citation impact, while also highlighting their limitations as tools for research impact assessment.

연구 동기 및 목표

  • 통계학에서 초기 단계 텍스트로부터 고도로 인용될 논문을 예측하는 도전을 동기 부여하고 형식화한다.
  • 출판 시점 정보(제목, 초록, 키워드, 연도, 장소)를 사용하여 고 영향력을 예측하는 텍스트 중심의 프롬프트 주도 프레임워크를 구축한다.
  • 다양한 '고도로 인용된' 정의(상위 퍼센트)에 따른 시간 간 예측 성능 및 모델 일반화를 평가한다.
  • 예측된 높은 영향력 논문과 연관된 주제적 신호(주제)를 식별한다.
  • 초기 단계 인용 영향력 평가를 돕기 위한 실용 도구(WeChat 미니 프로그램)를 제공한다.

제안 방법

  • LLMs에 대한 task를 방법론적 혁신, 장기 가치, 문제의 중요성을 평가하는 것으로 설정하는 구조화된 프롬프트 프레임워크를 개발한다.
  • 동시대 추세를 반영하기 위해 5년 창마다 시간적 배경 요약을 포함한다.
  • 판단을 보정하기 위해 여섯 개의 참고 예시(세 개는 highly cited, 세 개는 not)로 사용한다.
  • 출판 시점 데이터(제목, 초록, 키워드, 연도, 발행처)로 입력을 제한하여 포스트 퍼블리케이션 정보 누출을 피한다.
  • API(ChatGPT 4o mini, DeepSeek, Gemini)를 통해 다수의 LLM과 상호 작용하고 정확도, 진정/거짓 양성률, 효율성을 평가한다.
  • 논문이 높은 인용을 받을지 여부를 YES/NO 이진 예측으로 정의한다(상위 1%, 5%, 또는 10%).

실험 결과

연구 질문

  • RQ1LLMs가 초기 텍스트 및 맥락 정보를 통해 논문의 미래 인용 잠재력을 추론할 수 있는가?
  • RQ2구조화된 프롬프트를 사용하여 서로 다른 LLM들이 높은 인용 논문 예측에서 어떻게 비교되는가?
  • RQ3예측이 시간에 따라 일관되는가, 특정 주제나 방법론적 경향에 대한 주의를 드러내는가?
  • RQ4통계학에서 LLM 기반 영향력 예측의 실질적 한계와 편향은 무엇인가?
  • RQ5해당 접근법을 연구자와 기관이 사용할 수 있는 도구로 번역할 수 있는가?

주요 결과

  • LLMs는 초기 텍스트 데이터만을 사용하여 높은 인용을 받는 통계 논문을 예측하는 데 안정적이고 경쟁력 있는 성능을 보인다.
  • 예측 능력은 시간에 따라 일반화되어 출판 기간 전반에 걸쳐 유의미한 정보를 남긴다.
  • 예측된 높은 영향력 논문은 인과 추론과 딥 러닝과 같은 주제에서 군집하는 경향이 있다.
  • 실용적인 WeChat 미니 프로그램(Stat Highly Cited Papers)이 개발되어 초기 단계 영향력 평가를 촉진한다.
  • 본 연구는 LLM 기반 연구 영향력 평가의 가능성과 한계 모두를 강조한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.