[論文レビュー] Are Large Language Models able to Predict Highly Cited Papers? Evidence from Statistical Publications
本研究は、早期の文章情報だけを用いてLLMが将来の高インパクト統計論文を予測できるかを検証し、複数のモデルを比較し、プロンプト設計と時間的一般化の効果を明らかにします。
Predicting highly-cited papers is a long-standing challenge due to the complex interactions of research content, scholarly communities, and temporal dynamics. Recent advances in large language models (LLMs) raise the question of whether early-stage textual information can provide useful signals of long-term scientific impact. Focusing on statistical publications, we propose a flexible, text-centered framework that leverages LLMs and structured prompt design to predict highly cited papers. Specifically, we utilize information available at the time of publication, including titles, abstracts, keywords, and limited bibliographic metadata. Using a large corpus of statistical papers, we evaluate predictive performance across multiple publication periods and alternative definitions of highly cited papers. The proposed approach achieves stable and competitive performance relative to existing methods and demonstrates strong generalization over time. Textual analysis further reveals that papers predicted as highly cited concentrate on recurring topics such as causal inference and deep learning. To facilitate practical use of the proposed approach, we further develop a WeChat mini program, extit{Stat Highly Cited Papers}, which provides an accessible interface for early-stage citation impact assessment. Overall, our results provide empirical evidence that LLMs can capture meaningful early signals of long-term citation impact, while also highlighting their limitations as tools for research impact assessment.
研究の動機と目的
- 統計学における早期段階のテキストから高被引用論文を予測する課題の動機づけと正式化。
- 出版時情報(タイトル、要約、キーワード、年、会場)を用いて高インパクトを予測するテキスト中心の、プロンプト駆動フレームワークを構築。
- “高被引用”の定義(上位百分位)ごとに、時間を跨ぐ予測性能とモデルの一般化を評価。
- 予測される高インパクト論文に関連するテーマ的シグナル(トピック)を特定。
- 早期段階の引用影響評価を支援する実用的ツール(WeChatミニプログラム)を提供。
提案手法
- タスクをLLMに対して方法論革新、長期的価値、課題の重要性を評価する構造化プロンプトフレームワークとして定式化。
- contemporaneous trendsを反映させるため、各5年ウィンドウの時間的背景サマリーを組み込む。
- 判断を校正するために6つの基準例を用意(うち3つは高被引用、3つは非高被引用)。
- 後発情報のリークを避けるため、入力を出版時データ(タイトル、要約、キーワード、年、出版社)に限定。
- 複数のLLM(ChatGPT 4oミニ、DeepSeek、Gemini)とAPI経由で相互作用し、正確さ、真陽性/偽陽性率、効率を評価。
- タスクを高被引用であるか否かを二択のYES/NO予測として定義(上位1%、5%、または10%)。
実験結果
リサーチクエスチョン
- RQ1LLMは初期の文章および文脈情報から論文の将来の高被引用ポテンシャルを推測できるか。
- RQ2構造化プロンプトを用いて高被引用論文を予測する際、異なるLLMはどのように比較されるか。
- RQ3予測は時間とともに一貫性があるか、特定のトピックや方法論の動向に注意を向けているか。
- RQ4統計学におけるLLMベースの影響予測の実務的な限界とバイアスは何か。
- RQ5このアプローチを研究者や機関が実際に使えるツールへ翻訳できるか。
主な発見
- LLMsは早期のテキストデータのみを用いても、統計分野の高被引用論文を予測する能力が安定しており競争力があることを示す。
- 予測能力は時間を超えて一般化し、出版時期を超えて情報性を保つ。
- 予測された高インパクト論文は因果推論や深層学習などのトピックに clusterする傾向がある。
- 早期段階の影響評価を支援する実用的なWeChatミニプログラム(Stat Highly Cited Papers)を開発。
- 本研究はLLMベースの研究影響評価の可能性と制限の両方を示唆する。
より良い研究を、今すぐ始めましょう
論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。
クレジットカード登録不要
このレビューはAIが作成し、人間の編集者が確認しました。