[论文解读] Can social media provide early warning of retraction? Evidence from critical tweets identified by human annotation and large language models
本研究探讨了社交媒体,特别是批判性推文,是否可作为科学论文撤稿的早期预警系统。通过结合人工标注的批判性推文(撤稿论文中占8.3%,非撤稿论文中占1.5%)与大语言模型检测(GPT-4o mini、Gemini 2.0 Flash-Lite、Claude 3.5 Haiku),研究发现人工标注的信号显著优于仅使用大语言模型的结果,表明人机协作模式对于实现可扩展且准确的研究诚信监控至关重要。
Timely detection of problematic research is essential for safeguarding scientific integrity. To explore whether social media commentary can serve as an early indicator of potentially problematic articles, this study analysed 3,815 tweets referencing 604 retracted articles and 3,373 tweets referencing 668 comparable non-retracted articles. Tweets critical of the articles were identified through both human annotation and large language models (LLMs). Human annotation revealed that 8.3% of retracted articles were associated with at least one critical tweet prior to retraction, compared to only 1.5% of non-retracted articles, highlighting the potential of tweets as early warning signals of retraction. However, critical tweets identified by LLMs (GPT-4o mini, Gemini 2.0 Flash-Lite, and Claude 3.5 Haiku) only partially aligned with human annotation, suggesting that fully automated monitoring of post-publication discourse should be applied with caution. A human-AI collaborative approach may offer a more reliable and scalable alternative, with human expertise helping to filter out tweets critical of issues unrelated to the research integrity of the articles. Overall, this study provides insights into how social media signals, combined with generative AI technologies, may support efforts to strengthen research integrity.
研究动机与目标
- 评估社交媒体上针对科学论文的批判性言论是否可作为论文撤稿的早期预警信号。
- 比较人工标注的批判性推文与大语言模型(LLM)检测的推文,在识别后续被撤稿的论文方面的有效性。
- 评估大语言模型在检测与研究诚信相关的批评方面与人类判断的一致性与可靠性。
- 提出一种人机协作框架,以实现对研究诚信问题的后发表言论进行可扩展且准确的监控。
提出的方法
- 收集了涉及604篇撤稿论文的3,815条推文,以及涉及668篇非撤稿论文的3,373条推文。
- 通过人工标注识别出涉及研究诚信问题的批判性推文,建立了金标准数据集。
- 采用三种大语言模型——GPT-4o mini、Gemini 2.0 Flash-Lite 和 Claude 3.5 Haiku——对推文进行分类,判断其是否为批判性言论。
- 将大语言模型的预测结果与人工标注标签进行对比,以评估其一致性与可靠性。
- 计算精确率、召回率与F1分数,以评估大语言模型在检测与研究诚信相关批评方面的性能。
- 提出一种人机协作模型,由人工监督筛选并剔除大语言模型输出中与研究诚信无关的批评内容。
实验结果
研究问题
- RQ1社交媒体上的批判性推文能否作为后续被撤稿的科学论文的早期指标?
- RQ2与人工标注的标准相比,大语言模型在检测推文中与研究诚信相关的批评方面表现如何?
- RQ3人工标注的批判性推文与大语言模型生成的分类结果之间的一致性水平如何?
- RQ4在自动化监控后发表言论以识别研究诚信问题方面,大语言模型单独使用可信赖程度如何?
- RQ5人机协作模式在提升撤稿早期预警系统的可扩展性与准确性方面可发挥何种作用?
主要发现
- 在撤稿论文中,有8.3%与至少一条人工标注的批判性推文相关,而非撤稿论文中仅为1.5%。
- 人工标注数据集显示出强烈信号:在相当一部分案例中,批判性言论出现在撤稿之前。
- 大语言模型对批判性推文的检测与人工标注标签仅有部分一致性,表明其在完全自动化场景下可靠性有限。
- GPT-4o mini、Gemini 2.0 Flash-Lite 与 Claude 3.5 Haiku 在识别与研究诚信相关的批评时均表现出显著的假阳性和假阴性率。
- 研究结论认为,由于误分类风险,应谨慎使用完全自动化的大型语言模型监控。
- 建议采用人机协作模式,由人类专家对大语言模型的输出进行筛选与验证,以提升准确度与相关性。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。