Skip to main content
QUICK REVIEW

[论文解读] From Generalist to Specialist: A Survey of Large Language Models for Chemistry

Yang Han, Ziping Wan|arXiv (Cornell University)|Dec 28, 2024
Biomedical Text Mining and OntologiesBiochemistry, Genetics and Molecular Biology被引用 3
一句话总结

本综述系统性地回顾了专用于化学领域的大型语言模型(LLMs),重点关注领域特定知识的整合、多模态数据(2D/3D结构、光谱)以及工具使用能力。该综述提出了将通用型LLM转化为化学专用智能体的方法分类体系,评估了基准测试,并识别出推动化学科学发现的关键挑战与未来方向。

ABSTRACT

Large Language Models (LLMs) have significantly transformed our daily life and established a new paradigm in natural language processing (NLP). However, the predominant pretraining of LLMs on extensive web-based texts remains insufficient for advanced scientific discovery, particularly in chemistry. The scarcity of specialized chemistry data, coupled with the complexity of multi-modal data such as 2D graph, 3D structure and spectrum, present distinct challenges. Although several studies have reviewed Pretrained Language Models (PLMs) in chemistry, there is a conspicuous absence of a systematic survey specifically focused on chemistry-oriented LLMs. In this paper, we outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs, we also conceptualize chemistry LLMs as agents using chemistry tools and investigate their potential to accelerate scientific research. Additionally, we conclude the existing benchmarks to evaluate chemistry ability of LLMs. Finally, we critically examine the current challenges and identify promising directions for future research. Through this comprehensive survey, we aim to assist researchers in staying at the forefront of developments in chemistry LLMs and to inspire innovative applications in the field.

研究动机与目标

  • 为解决通用型LLM在化学领域中的局限性,识别在领域知识、多模态数据和工具集成方面面临的关键挑战。
  • 提供一种全面的分类体系,用于描述通过预训练、微调及工具增强推理,将通用LLM转化为专用化学LLM的方法。
  • 评估现有基准测试与数据集,以评估化学专用LLM在文本、结构和光谱模态下的能力。
  • 探讨LLM作为自主智能体,利用化学工具(如分子编辑器、模拟引擎)实现科学工作流加速的潜力。
  • 识别在数据稀缺性、模型对齐性及多模态推理方面存在的开放挑战与未来研究方向。

提出的方法

  • 将化学LLM的适应方法归纳为三大核心挑战:领域知识整合、多模态数据处理(1D序列、2D图、3D结构、光谱)以及工具使用。
  • 回顾预训练、监督微调(SFT)和基于人类反馈的强化学习(RLHF)作为关键适应策略,其中SFT进一步细分为任务特定与多任务设置。
  • 分析处理2D分子图(如InstructMol、MolTC)、3D结构(如3D-MoLM)和图像(如ChemVLM、MM-RCR)的多模态LLM。
  • 考察利用外部工具(如分子生成器、反应预测器、模拟引擎)的工具增强型LLM(如Coscientist、ChemCrow、LLaMP)。
  • 提出一种将化学LLM建模为智能体的框架,支持规划、知识检索与结构化工具执行动作。
  • 综述14个以上基准测试(如ChemLLMBench、MassSpecGym、ScholarChemQA),涵盖模态(文本、光谱、图像)与任务类型(选择题、直接回答)。
Figure 1: Three common errors in general LLMs arising from the key challenges.
Figure 1: Three common errors in general LLMs arising from the key challenges.

实验结果

研究问题

  • RQ1如何通过领域知识整合,有效提升通用型LLM处理复杂化学专用任务的能力?
  • RQ2在面向化学的LLM中,最有效的多模态数据(如2D/3D分子结构、光谱、图像)整合方法是什么?
  • RQ3LLM在多大程度上可通过利用外部化学工具实现自主智能体功能,以支持科学推理与合成规划?
  • RQ4现有基准测试在评估化学LLM时存在哪些当前局限性与瓶颈?
  • RQ5在化学LLM领域,哪些未来研究方向最具前景,可推动该技术的最前沿发展?

主要发现

  • 通用型LLM因领域知识不足,在反应预测、性质计算和命名方面表现不佳,导致错误频发。
  • 多模态LLM如3D-MoLM和ChemVLM在3D结构与图像任务上表现更优,但仍受限于数据稀缺性与标注质量。
  • 基于SFT的模型如LlaSMol和ChemDFM在分子性质预测与反应生成任务中表现强劲,相比零样本提示显著提升了准确率。
  • 工具增强型LLM如ChemCrow与LLaMP通过集成外部工具,在逆合成与试剂预测等复杂任务中实现了显著性能提升。
  • MassSpecGym与ScholarChemQA等基准测试揭示,LLM在光谱解析与领域特定推理方面仍存在困难,尤其在分布外设置下表现更差。
  • 本综述识别出一个关键缺口:缺乏开放、大规模、多模态数据集,特别是真实分子图像与反应机理图,严重限制了模型的泛化能力。
Figure 3: For example, the compound $C_{8}H_{11}NO$ can be represented across various modalities. 1D sequeues include SMILES, IUPAC name and so on. Molecular structure consist of 2D graphs and 3D structures, 2D graphs encompass three matrices: atomic features, atom connection, chemical bonds feature
Figure 3: For example, the compound $C_{8}H_{11}NO$ can be represented across various modalities. 1D sequeues include SMILES, IUPAC name and so on. Molecular structure consist of 2D graphs and 3D structures, 2D graphs encompass three matrices: atomic features, atom connection, chemical bonds feature

更好的研究,从现在开始

从阅读论文到最终审阅,大幅缩短您的研究时间。

无需绑定信用卡

本解读由 AI 生成,并经人工编辑审核。