[论文解读] ChemEval: A Comprehensive Multi-Level Chemical Evaluation for Large Language Models
ChemEval 是一个全面的、多层级的基准测试,用于评估大语言模型(LLMs)在化学领域的表现,涵盖 42 项任务、12 个维度和 4 个逐步提升的复杂度级别。该基准在零样本和少样本设置下评估了 12 种主流 LLM,结果表明:在高级化学推理任务中,专业 LLM 表现优于通用模型,而通用模型在指令遵循和文献理解方面表现更优。
There is a growing interest in the role that LLMs play in chemistry which lead to an increased focus on the development of LLMs benchmarks tailored to chemical domains to assess the performance of LLMs across a spectrum of chemical tasks varying in type and complexity. However, existing benchmarks in this domain fail to adequately meet the specific requirements of chemical research professionals. To this end, we propose extbf{ extit{ChemEval}}, which provides a comprehensive assessment of the capabilities of LLMs across a wide range of chemical domain tasks. Specifically, ChemEval identified 4 crucial progressive levels in chemistry, assessing 12 dimensions of LLMs across 42 distinct chemical tasks which are informed by open-source data and the data meticulously crafted by chemical experts, ensuring that the tasks have practical value and can effectively evaluate the capabilities of LLMs. In the experiment, we evaluate 12 mainstream LLMs on ChemEval under zero-shot and few-shot learning contexts, which included carefully selected demonstration examples and carefully designed prompts. The results show that while general LLMs like GPT-4 and Claude-3.5 excel in literature understanding and instruction following, they fall short in tasks demanding advanced chemical knowledge. Conversely, specialized LLMs exhibit enhanced chemical competencies, albeit with reduced literary comprehension. This suggests that LLMs have significant potential for enhancement when tackling sophisticated tasks in the field of chemistry. We believe our work will facilitate the exploration of their potential to drive progress in chemistry. Our benchmark and analysis will be available at {\color{blue} \url{https://github.com/USTC-StarTeam/ChemEval}}.
研究动机与目标
- 为解决当前缺乏全面、领域特定的基准测试来评估 LLM 在化学领域的表现这一问题。
- 识别并评估从基础知识到高级推理的 4 个逐步提升的化学推理层级。
- 通过专家精心设计的数据和提示构建基准,以确保实际相关性和评估的严谨性。
- 在零样本和少样本设置下,对多样化的化学任务中的通用和专业 LLM 进行评估。
- 为支持真实世界化学研究提供关于 LLM 优势与局限性的可操作洞察。
提出的方法
- 设计一个包含 4 个逐步提升层级的多层级框架:基础知识、推理、推断和高级问题解决。
- 定义 12 个评估维度,涵盖化学知识、反应预测、产率估算和机理分析。
- 使用开源数据和专家验证的示例,精心整理出 42 项独立的化学任务,以确保科学准确性。
- 开发针对每项任务的提示,包含系统指令和少样本示例,以评估零样本和少样本的泛化能力。
- 整合专家验证的 SMILES 字符串和反应条件,以确保预测结果的化学正确性。
- 在标准化提示协议下,对 12 种最先进的 LLM(包括 GPT-4 和 Claude-3.5)在所有任务上进行评估。
实验结果
研究问题
- RQ1与专业化学 LLM 相比,通用 LLM 在复杂化学推理任务中的表现如何?
- RQ2LLM 仅通过 SMILES 输入,能在多大程度上准确预测反应产物、产率和活化能?
- RQ3LLM 能否从反应物和产物中推导出合理的反应中间体?这些推导结果的可靠性如何?
- RQ4少样本提示对需要深厚领域知识的化学任务中 LLM 表现有何影响?
- RQ5当前 LLM 在支持高级化学研究(尤其是合成与机理预测)方面存在哪些关键能力缺口?
主要发现
- 通用 LLM(如 GPT-4 和 Claude-3.5)在文献理解与指令遵循方面表现优异,但在高级化学推理任务中表现吃力。
- 专业 LLM 在复杂任务(如反应机理推导和产率预测)中表现出显著提升的性能。
- 反应结果预测(PPred)在专家精心设计的示例上达到了高准确率,基准中报告了正确的 SMILES 输出。
- 产率预测(YPred)显示,模型能够以合理可靠性区分高产率(≥70%)与低产率反应。
- 活化能预测(RatePred)表明,模型能够以可接受的精度估算范围(例如,乙醇脱水反应的活化能范围为 180–370 kJ/mol)。
- 中间体推导(IMDer)表明,模型能够正确识别常见机理中的关键物种(如碳正离子),但在复杂案例中存在显著错误率。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。