[Paper Review] From Words to Molecules: A Survey of Large Language Models in Chemistry
This survey categorizes how large language models (LLMs) are adapted for chemistry, detailing molecule representations, tokenization, pretraining objectives, and application paradigms. It also outlines future research directions.
In recent years, Large Language Models (LLMs) have achieved significant success in natural language processing (NLP) and various interdisciplinary areas. However, applying LLMs to chemistry is a complex task that requires specialized domain knowledge. This paper provides a thorough exploration of the nuanced methodologies employed in integrating LLMs into the field of chemistry, delving into the complexities and innovations at this interdisciplinary juncture. Specifically, our analysis begins with examining how molecular information is fed into LLMs through various representation and tokenization methods. We then categorize chemical LLMs into three distinct groups based on the domain and modality of their input data, and discuss approaches for integrating these inputs for LLMs. Furthermore, this paper delves into the pretraining objectives with adaptations to chemical LLMs. After that, we explore the diverse applications of LLMs in chemistry, including novel paradigms for their application in chemistry tasks. Finally, we identify promising research directions, including further integration with chemical knowledge, advancements in continual learning, and improvements in model interpretability, paving the way for groundbreaking developments in the field.
Motivation & Objective
- Systematically review how molecular information is tokenized and represented for LLMs.
- Provide a taxonomy of chemical LLMs based on input domain and modality (single-domain, multi-domain, multi-modal).
- Analyze pretraining objectives and adaptation strategies for chemical data.
- Explore diverse chemistry applications enabled by LLMs and identify open research directions.
Proposed method
- Classifies molecular representations (fingerprints, SMILES/SELFIES, InChI, graph-based) and tokenization levels (character-, atom-, motif-level).
- Presents a taxonomy of pretraining data domains (single-domain, multi-domain, multi-modal) and integration strategies.
- Reviews three core pretraining objectives for chemical LLMs: Masked Language Modeling (MLM), Molecule Property Prediction (MPP), and Autoregressive Token Generation (ATG) with chemistry-specific tasks.
- Discusses cross-modal objectives such as cross-modal contrastive learning (XMC) and alignment across modalities.
- Summarizes representative architectures, datasets, and training approaches in a comprehensive table of methods.
- Outlines applications and future directions including continual learning and interpretability.
Experimental results
Research questions
- RQ1How are molecular sequences tokenized and represented for LLMs in chemistry?
- RQ2What taxonomy best captures chemical LLMs by input domain and modality?
- RQ3What pretraining objectives are used and how are they adapted to chemical data?
- RQ4What are the primary applications and paradigms enabled by chemical LLMs?
- RQ5What future directions can advance integration of chemical knowledge, continual learning, and interpretability?
Key findings
- Molecule representations include fingerprints, SMILES/SELFIES, InChI, and graph-based forms with varying granularity.
- Tokenization schemes span character-, atom-, and motif-level approaches, with data-driven and chemistry-driven methods.
- Chemical LLMs are organized into single-domain, multi-domain, and multi-modal taxonomies depending on input data and modalities.
- MLM, MPP, and ATG are the core pretraining objectives, with MPP providing strong representation learning signals and ATG enabling task alignment.
- Cross-modal learning and representation alignment are used to fuse text, graphs, fingerprints, and images, though domain-specific nuances remain a challenge.
- Applications encompass chatbots, in-context learning, and representation learning for downstream tasks like property prediction, reaction prediction, and molecule generation.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.