[论文解读] Journal Impact Factor and Peer Review Thoroughness and Helpfulness: A Supervised Machine Learning Study
本研究利用监督式机器学习分析了来自1,644本医学与生命科学期刊的187,240条同行评审语句,并将其与期刊影响因子(JIF)分位数关联。研究发现,高JIF期刊的评审更具方法论细节,但提供的帮助性反馈更少,建议与实例也更少,表明尽管内容分布存在微小差异,JIF作为个体稿件评审质量的预测指标仍较弱。
The journal impact factor (JIF) is often equated with journal quality and the quality of the peer review of the papers submitted to the journal. We examined the association between the content of peer review and JIF by analysing 10,000 peer review reports submitted to 1,644 medical and life sciences journals. Two researchers hand-coded a random sample of 2,000 sentences. We then trained machine learning models to classify all 187,240 sentences as contributing or not contributing to content categories. We examined the association between ten groups of journals defined by JIF deciles and the content of peer reviews using linear mixed-effects models, adjusting for the length of the review. The JIF ranged from 0.21 to 74.70. The length of peer reviews increased from the lowest (median number of words 185) to the JIF group (387 words). The proportion of sentences allocated to different content categories varied widely, even within JIF groups. For thoroughness, sentences on 'Materials and Methods' were more common in the highest JIF journals than in the lowest JIF group (difference of 7.8 percentage points; 95% CI 4.9 to 10.7%). The trend for 'Presentation and Reporting' went in the opposite direction, with the highest JIF journals giving less emphasis to such content (difference -8.9%; 95% CI -11.3 to -6.5%). For helpfulness, reviews for higher JIF journals devoted less attention to 'Suggestion and Solution' and provided fewer Examples than lower impact factor journals. No, or only small differences were evident for other content categories. In conclusion, peer review in journals with higher JIF tends to be more thorough in discussing the methods used but less helpful in terms of suggesting solutions and providing examples. Differences were modest and variability high, indicating that the JIF is a bad predictor for the quality of peer review of an individual manuscript.
研究动机与目标
- 探究期刊影响因子(JIF)是否与医学与生命科学期刊同行评审的全面性与帮助性相关。
- 利用大规模评审报告数据集,通过监督式机器学习对同行评审内容进行标准化分类。
- 评估高JIF期刊相较于低JIF期刊是否提供更全面或更具建设性的反馈。
- 在评审内容存在差异的前提下,评估JIF对个体稿件评审质量的预测能力。
提出的方法
- 由两名研究人员对10,000份同行评审报告中随机抽取的2,000条语句进行人工编码,划分为十个内容类别。
- 训练监督式机器学习模型,以高精度将剩余的187,240条语句分类至上述内容类别。
- 采用线性混合效应模型分析JIF分位数与内容分布之间的关联,并调整了评审长度的影响。
- 内容类别包括“材料与方法”、“展示与报告”、“建议与解决方案”、“实例”等,重点关注全面性与帮助性指标。
- 使用JIF分位数将期刊划分为10个组别(10个分位数),以实现不同影响力水平期刊间评审内容的比较。
实验结果
研究问题
- RQ1期刊影响因子与同行评审全面性之间是否存在显著关联,特别是在讨论方法与报告方面?
- RQ2高影响因子期刊是否提供更多具有帮助性的反馈,如改进建议或实例说明?
- RQ3不同影响力水平期刊的同行评审内容分布如何变化?这些模式的一致性如何?
- RQ4在个体稿件评审质量的预测方面,期刊影响因子的预测效力有多大?
主要发现
- 最高JIF组期刊的评审中,‘材料与方法’相关内容占比比最低JIF组高出7.8个百分点(95%置信区间:4.9至10.7个百分点)。
- 最高JIF组期刊的评审中,对‘展示与报告’的关注度比最低JIF组低8.9个百分点(95%置信区间:-11.3至-6.5个百分点)。
- 高JIF期刊提供的‘建议与解决方案’类评论及‘实例’类内容显著少于低JIF期刊,表明其帮助性较低。
- 即使在同一JIF组内,多数内容类别的句子比例也存在广泛差异,表明组内变异程度较高。
- 在‘清晰性与语言’或‘伦理考量’等其他内容类别中未观察到显著差异。
- 总体而言,由于效应量微小且评审内容变异程度高,本研究得出结论:JIF对个体稿件评审质量的预测能力较差。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。