[Paper Review] Usefulness of LLMs as an Author Checklist Assistant for Scientific Papers: NeurIPS'24 Experiment
This study evaluates a large language model (LLM)-powered checklist assistant deployed at NeurIPS 2024 to help authors verify compliance with submission standards. The assistant provided specific, actionable feedback on 234 papers, with over 70% of users finding it useful and revising their work based on its input, though issues like inaccuracy and excessive strictness were reported; the tool showed promise in improving submission quality despite vulnerabilities to manipulation.
Large language models (LLMs) represent a promising, but controversial, tool in aiding scientific peer review. This study evaluates the usefulness of LLMs in a conference setting as a tool for vetting paper submissions against submission standards. We conduct an experiment at the 2024 Neural Information Processing Systems (NeurIPS) conference, where 234 papers were voluntarily submitted to an "LLM-based Checklist Assistant." This assistant validates whether papers adhere to the author checklist used by NeurIPS, which includes questions to ensure compliance with research and manuscript preparation standards. Evaluation of the assistant by NeurIPS paper authors suggests that the LLM-based assistant was generally helpful in verifying checklist completion. In post-usage surveys, over 70% of authors found the assistant useful, and 70% indicate that they would revise their papers or checklist responses based on its feedback. While causal attribution to the assistant is not definitive, qualitative evidence suggests that the LLM contributed to improving some submissions. Survey responses and analysis of re-submissions indicate that authors made substantive revisions to their submissions in response to specific feedback from the LLM. The experiment also highlights common issues with LLMs: inaccuracy (20/52) and excessive strictness (14/52) were the most frequent issues flagged by authors. We also conduct experiments to understand potential gaming of the system, which reveal that the assistant could be manipulated to enhance scores through fabricated justifications, highlighting potential vulnerabilities of automated review tools.
Motivation & Objective
- To evaluate the usefulness of an LLM-based checklist assistant in helping authors meet NeurIPS submission standards.
- To assess author perceptions and behavioral changes following interaction with the LLM assistant.
- To identify common failure modes and vulnerabilities of LLMs in a real-world academic checklist verification setting.
- To explore whether LLM feedback leads to measurable improvements in paper quality and checklist completeness.
Proposed method
- Deployed an LLM-powered checklist assistant that evaluates authors’ responses to the NeurIPS 2024 Paper Checklist, a 15-question yes/no checklist on reproducibility, ethics, and transparency.
- The assistant analyzes each checklist response and generates detailed, specific feedback on gaps or improvements needed, such as referencing specific sections or expanding ethical impact discussions.
- Conducted pre- and post-usage surveys with 539 and 78 respondents, respectively, to assess user expectations and experiences.
- Collected and analyzed 234 submissions to the assistant, including 40 re-submissions, to track changes in checklist justification length and content.
- Conducted adversarial experiments to test whether the system could be manipulated via fabricated justifications to improve scores.
Experimental results
Research questions
- RQ1To what extent do authors find an LLM-based checklist assistant useful in preparing their NeurIPS submissions?
- RQ2Does feedback from the LLM assistant lead to measurable improvements in checklist completeness and manuscript quality?
- RQ3What are the most common failure modes of LLMs when used for checklist verification, such as inaccuracy or excessive strictness?
- RQ4Can the LLM-based checklist system be easily manipulated through deceptive justifications?
Key findings
- Over 70% of authors found the LLM-based checklist assistant useful and reported they would revise their papers or checklist responses based on its feedback.
- Authors made substantive revisions in response to the assistant, including significantly increasing the length and specificity of their checklist justifications between re-submissions.
- The LLM provided an average of 4–6 distinct, specific feedback points per checklist question, demonstrating granular analysis of submission content.
- The most frequently cited issues were inaccuracy (20 out of 52 respondents) and excessive strictness (14 out of 52 respondents).
- The system was vulnerable to manipulation, with fabricated justifications successfully increasing scores in adversarial testing, highlighting risks in automated review tools.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.