[论文解读] From Words to Molecules: A Survey of Large Language Models in Chemistry
本综述分类了大语言模型(LLMs)如何被改编用于化学,详述分子表示、标记化、预训练目标以及应用范式。它还概述了未来的研究方向。
In recent years, Large Language Models (LLMs) have achieved significant success in natural language processing (NLP) and various interdisciplinary areas. However, applying LLMs to chemistry is a complex task that requires specialized domain knowledge. This paper provides a thorough exploration of the nuanced methodologies employed in integrating LLMs into the field of chemistry, delving into the complexities and innovations at this interdisciplinary juncture. Specifically, our analysis begins with examining how molecular information is fed into LLMs through various representation and tokenization methods. We then categorize chemical LLMs into three distinct groups based on the domain and modality of their input data, and discuss approaches for integrating these inputs for LLMs. Furthermore, this paper delves into the pretraining objectives with adaptations to chemical LLMs. After that, we explore the diverse applications of LLMs in chemistry, including novel paradigms for their application in chemistry tasks. Finally, we identify promising research directions, including further integration with chemical knowledge, advancements in continual learning, and improvements in model interpretability, paving the way for groundbreaking developments in the field.
研究动机与目标
- 系统性地评估分子信息如何被标记化并用于 LLMs。
- 基于输入域和模态为化学 LLMs 提供一个分类法(单域、多域、多模态)。
- 分析化学数据的预训练目标和适应策略。
- 探索 LLMs 支持的多样化化学应用并识别未解的研究方向。
提出的方法
- 对分子表示进行分类(指纹、SMILES/SELFIES、InChI、基于图的)以及标记化层级(字符级、原子级、基元级)。
- 提出预训练数据域的分类法(单域、多域、跨模态)及整合策略。
- 回顾化学 LLMs 的三个核心预训练目标:Masked Language Modeling (MLM)、Molecule Property Prediction (MPP)、Autoregressive Token Generation (ATG) 及化学特定任务。
- 讨论跨模态目标,如跨模态对比学习(XMC)及模态之间的对齐。
- 在方法的综合表中总结具有代表性的架构、数据集和训练方法。
- 概述应用与未来方向,包括持续学习和可解释性。
实验结果
研究问题
- RQ1分子序列在化学中如何被标记化并表示给 LLMs?
- RQ2哪种分类法最能通过输入域和模态来描述化学 LLMs?
- RQ3使用了哪些预训练目标,它们如何适应化学数据?
- RQ4化学 LLMs 支持的主要应用和范式有哪些?
- RQ5哪些未来方向能推动化学知识整合、持续学习和可解释性?
主要发现
- 分子表征包括指纹、SMILES/SELFIES、InChI,以及基于图的形式,具有不同的粒度。
- 标记化方案涵盖字符级、原子级和基元级方法,具备数据驱动与化学驱动两种途径。
- 化学 LLMs 根据输入数据和模态的不同,组织为单域、多域和多模态分类法。
- MLM、MPP 与 ATG 是核心预训练目标,MPP 提供强大的表征学习信号,ATG 使任务对齐成为可能。
- 跨模态学习与表示对齐被用于将文本、图、指纹和图像融合,尽管领域特定的细微差异仍是挑战。
- 应用包括聊天机器人、上下文学习,以及用于下游任务如性质预测、反应预测和分子生成的表示学习。
更好的研究,从现在开始
从阅读论文到最终审阅,大幅缩短您的研究时间。
无需绑定信用卡
本解读由 AI 生成,并经人工编辑审核。