Skip to main content
QUICK REVIEW

[論文レビュー] Journal Impact Factor and Peer Review Thoroughness and Helpfulness: A Supervised Machine Learning Study

Anna Severin, Michaela Strinzel|arXiv (Cornell University)|Jul 20, 2022
Meta-analysis and systematic reviews被引用数 7
ひとこと要約

本研究では、1,644 冊の医学・生命科学分野の学術雑誌から得られた 187,240 文の査読文を、ジャーナル・インパクト・ファクター(JIF)のデシルと関連付けて分析するための教師あり機械学習が用いられた。その結果、JIF が高い雑誌では、より方法論的に詳細な査読がなされるが、助言や例示の数が少なく、支援的でないフィードバックが増える傾向にあり、内容分布にわずかな差があるものの、JIF は個々の査読の質を弱い指標にとどめることが示された。

ABSTRACT

The journal impact factor (JIF) is often equated with journal quality and the quality of the peer review of the papers submitted to the journal. We examined the association between the content of peer review and JIF by analysing 10,000 peer review reports submitted to 1,644 medical and life sciences journals. Two researchers hand-coded a random sample of 2,000 sentences. We then trained machine learning models to classify all 187,240 sentences as contributing or not contributing to content categories. We examined the association between ten groups of journals defined by JIF deciles and the content of peer reviews using linear mixed-effects models, adjusting for the length of the review. The JIF ranged from 0.21 to 74.70. The length of peer reviews increased from the lowest (median number of words 185) to the JIF group (387 words). The proportion of sentences allocated to different content categories varied widely, even within JIF groups. For thoroughness, sentences on 'Materials and Methods' were more common in the highest JIF journals than in the lowest JIF group (difference of 7.8 percentage points; 95% CI 4.9 to 10.7%). The trend for 'Presentation and Reporting' went in the opposite direction, with the highest JIF journals giving less emphasis to such content (difference -8.9%; 95% CI -11.3 to -6.5%). For helpfulness, reviews for higher JIF journals devoted less attention to 'Suggestion and Solution' and provided fewer Examples than lower impact factor journals. No, or only small differences were evident for other content categories. In conclusion, peer review in journals with higher JIF tends to be more thorough in discussing the methods used but less helpful in terms of suggesting solutions and providing examples. Differences were modest and variability high, indicating that the JIF is a bad predictor for the quality of peer review of an individual manuscript.

研究の動機と目的

  • 医学・生命科学分野の学術雑誌におけるジャーナル・インパクト・ファクター(JIF)と査読の包括性および支援性の相関関係を調査すること。
  • 大規模な査読レポートのデータセットを用いて教師あり機械学習により、査読内容を標準化されたカテゴリに分類すること。
  • 高い JIF の雑誌が、低い JIF の雑誌よりも、より包括的またはより建設的なフィードバックを提供しているかどうかを評価すること。
  • 査読内容のばらつきを踏まえ、JIF が個々の査読の質を予測する力はどの程度かを評価すること。

提案手法

  • 2 名の研究者が、10,000 件の査読レポートからランダムに抽出した 2,000 文を、10 のコンテンツカテゴリに手動でコード化した。
  • 残りの 187,240 文を、高い正確性でこれらのコンテンツカテゴリに分類するための教師あり機械学習モデルを訓練した。
  • JIF デシルとコンテンツ分布の関連を、査読の長さを補正したうえで分析するために、線形混合効果モデルを用いた。
  • コンテンツカテゴリには「材料と方法」「報告と提示」「助言と解決策」「例示」などがあり、主に包括性と支援性の指標に注目した。
  • JIF デシル(10 グループ)を用いて雑誌をグループ化し、インパクトレベルごとの査読内容を比較可能にした。

実験結果

リサーチクエスチョン

  • RQ1ジャーナル・インパクト・ファクターと査読の包括性、特に方法論や報告に関する議論の有無に有意な関連があるか。
  • RQ2高いインパクト・ファクターの雑誌は、改善のための助言や具体的な例示を、より多く提供しているか。
  • RQ3異なるインパクトレベルの雑誌における査読コンテンツの分布はどのように異なるのか。また、これらのパターンは一貫性があるか。
  • RQ4JIF は、個々の原稿の査読の質をどの程度まで予測できるか。

主な発見

  • 最高 JIF グループの査読では、最低 JIF グループと比較して「材料と方法」に関する内容が 7.8 パcent 点多い(95% CI 4.9 ~ 10.7%)。
  • 最高 JIF グループの査読では、最低 JIF グループと比較して「報告と提示」に割り当てられる注意が 8.9 パcent 点少ない(95% CI -11.3 ~ -6.5%)。
  • 高い JIF の雑誌では、低い JIF の雑誌と比較して「助言と解決策」のコメントが著しく少なく、例示の数も少ないため、支援性が低いことが示された。
  • 同じ JIF グループ内でも、ほとんどのコンテンツカテゴリにおける文の割合に広範なばらつきが認められ、グループ内での変動が大きいことが示された。
  • 「明確さと言語」や「倫理的配慮」などの他のコンテンツカテゴリには顕著な差は認められなかった。
  • 全体として、本研究は、JIF が個々の査読の質を予測するのに不適切であると結論づけた。これは、効果量が小さく、査読コンテンツのばらつきが大きいことが要因である。

より良い研究を、今すぐ始めましょう

論文の読解から最終レビューまで、研究時間を劇的に削減しましょう。

クレジットカード登録不要

このレビューはAIが作成し、人間の編集者が確認しました。