Skip to main content
QUICK REVIEW

[论文解读] Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment

Alexander Goldberg, Ihsan Ullah|arXiv (Cornell University)|Nov 5, 2024
Artificial Intelligence in Healthcare and EducationMedicine被引用 3
一句话总结

本研究评估了一款在 NeurIPS 2024 上部署的大型语言模型(LLM)驱动的检查清单助手,旨在帮助作者验证其投稿是否符合提交标准。该助手对 234 篇论文提供了具体且可操作的反馈,超过 70% 的用户认为其有用,并根据反馈修改了稿件,尽管也报告了准确性不足和过度严格等问题;该工具在提升投稿质量方面展现出潜力,但存在被操纵的漏洞。

ABSTRACT

Large language models (LLMs) represent a promising, but controversial, tool in aiding scientific peer review. This study evaluates the usefulness of LLMs in a conference setting as a tool for vetting paper submissions against submission standards. We conduct an experiment at the 2024 Neural Information Processing Systems (NeurIPS) conference, where 234 papers were voluntarily submitted to an "LLM-based Checklist Assistant." This assistant validates whether papers adhere to the author checklist used by NeurIPS, which includes questions to ensure compliance with research and manuscript preparation standards. Evaluation of the assistant by NeurIPS paper authors suggests that the LLM-based assistant was generally helpful in verifying checklist completion. In post-usage surveys, over 70% of authors found the assistant useful, and 70% indicate that they would revise their papers or checklist responses based on its feedback. While causal attribution to the assistant is not definitive, qualitative evidence suggests that the LLM contributed to improving some submissions. Survey responses and analysis of re-submissions indicate that authors made substantive revisions to their submissions in response to specific feedback from the LLM. The experiment also highlights common issues with LLMs: inaccuracy (20/52) and excessive strictness (14/52) were the most frequent issues flagged by authors. We also conduct experiments to understand potential gaming of the system, which reveal that the assistant could be manipulated to enhance scores through fabricated justifications, highlighting potential vulnerabilities of automated review tools.

研究动机与目标

  • 评估基于 LLM 的检查清单助手在帮助作者满足 NeurIPS 投稿标准方面的实用性。
  • 评估作者在与 LLM 助手互动后的感知与行为变化。
  • 识别 LLM 在真实世界学术检查清单验证场景中常见的故障模式与漏洞。
  • 探索 LLM 反馈是否能带来论文质量与检查清单完整性的可测量提升。

提出的方法

  • 部署了一款基于 LLM 的检查清单助手,用于评估作者对 NeurIPS 2024 论文检查清单(一项关于可复现性、伦理与透明度的 15 个是/否问题的清单)的回答。
  • 该助手分析每个检查清单的回答,并生成关于缺失内容或改进空间的详细、具体反馈,例如引用特定章节或扩展伦理影响讨论。
  • 通过前后两次调查(分别有 539 名和 78 名受访者)评估用户期望与使用体验。
  • 收集并分析了 234 份提交至该助手的稿件,包括 40 次重新提交,以追踪检查清单理由长度与内容的变化。
  • 开展对抗性实验,测试通过伪造理由是否可操纵系统以提高评分。

实验结果

研究问题

  • RQ1作者在准备 NeurIPS 投稿时,对基于 LLM 的检查清单助手的实用性认可程度如何?
  • RQ2LLM 助手的反馈是否带来了检查清单完整性和稿件质量的可测量提升?
  • RQ3LLM 在检查清单验证中常见的故障模式是什么,例如不准确或过度严格?
  • RQ4是否可通过欺骗性理由轻易操纵基于 LLM 的检查清单系统?

主要发现

  • 超过 70% 的作者认为基于 LLM 的检查清单助手有用,并报告会根据其反馈修改论文或检查清单回答。
  • 作者根据助手反馈进行了实质性修改,包括在重新提交时显著增加了检查清单理由的长度与具体性。
  • LLM 平均每道检查清单问题提供 4 至 6 个不同的具体反馈点,展示了对投稿内容的细致分析能力。
  • 最常见的问题为不准确(52 名受访者中有 20 人提及)和过度严格(52 名受访者中有 14 人提及)。
  • 该系统易受操纵,在对抗性测试中,伪造理由成功提高了评分,凸显了自动化评审工具的风险。

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。