[Paper Review] From Generalist to Specialist: A Survey of Large Language Models for Chemistry
This survey systematically reviews large language models (LLMs) specialized for chemistry, focusing on integrating domain-specific knowledge, multi-modal data (2D/3D structures, spectra), and tool-use capabilities. It proposes a taxonomy of methods for adapting generalist LLMs into chemistry-specific agents, evaluates benchmarks, and identifies key challenges and future directions for advancing scientific discovery in chemistry.
Large Language Models (LLMs) have significantly transformed our daily life and established a new paradigm in natural language processing (NLP). However, the predominant pretraining of LLMs on extensive web-based texts remains insufficient for advanced scientific discovery, particularly in chemistry. The scarcity of specialized chemistry data, coupled with the complexity of multi-modal data such as 2D graph, 3D structure and spectrum, present distinct challenges. Although several studies have reviewed Pretrained Language Models (PLMs) in chemistry, there is a conspicuous absence of a systematic survey specifically focused on chemistry-oriented LLMs. In this paper, we outline methodologies for incorporating domain-specific chemistry knowledge and multi-modal information into LLMs, we also conceptualize chemistry LLMs as agents using chemistry tools and investigate their potential to accelerate scientific research. Additionally, we conclude the existing benchmarks to evaluate chemistry ability of LLMs. Finally, we critically examine the current challenges and identify promising directions for future research. Through this comprehensive survey, we aim to assist researchers in staying at the forefront of developments in chemistry LLMs and to inspire innovative applications in the field.
Motivation & Objective
- To address the limitations of generalist LLMs in chemistry by identifying key challenges in domain knowledge, multi-modal data, and tool integration.
- To provide a comprehensive taxonomy of approaches for transforming general LLMs into specialized chemistry LLMs through pre-training, fine-tuning, and tool-augmented reasoning.
- To evaluate existing benchmarks and datasets for assessing chemistry-specific LLM capabilities, including text, structure, and spectral modalities.
- To explore the role of LLMs as autonomous agents using chemistry tools (e.g., molecular editors, simulation engines) to accelerate scientific workflows.
- To identify open challenges and future research directions in data scarcity, model alignment, and multi-modal reasoning for chemistry applications.
Proposed method
- Categorizes chemistry LLM adaptation into three core challenges: domain knowledge integration, multi-modal data handling (1D sequences, 2D graphs, 3D structures, spectra), and tool usage.
- Reviews pre-training, supervised fine-tuning (SFT), and reinforcement learning with human feedback (RLHF) as key adaptation strategies, with SFT further divided into task-specific and multi-task settings.
- Analyzes multi-modal LLMs that process 2D molecular graphs (e.g., InstructMol, MolTC), 3D structures (e.g., 3D-MoLM), and images (e.g., ChemVLM, MM-RCR).
- Examines tool-augmented LLMs that use external tools such as molecular generators, reaction predictors, and simulation engines (e.g., Coscientist, ChemCrow, LLaMP).
- Proposes a framework for modeling chemistry LLMs as agents that plan, retrieve knowledge, and execute actions using structured tools.
- Reviews 14+ benchmarks (e.g., ChemLLMBench, MassSpecGym, ScholarChemQA) across modalities (text, spectra, images) and tasks (MCQ, direct answer).

Experimental results
Research questions
- RQ1How can generalist LLMs be effectively adapted to handle complex chemistry-specific tasks through domain knowledge integration?
- RQ2What are the most effective methods for incorporating multi-modal data (e.g., 2D/3D molecular structures, spectra, images) into chemistry-oriented LLMs?
- RQ3To what extent can LLMs function as autonomous agents by leveraging external chemistry tools for scientific reasoning and synthesis planning?
- RQ4What are the current limitations and bottlenecks in evaluating chemistry LLMs using existing benchmarks?
- RQ5What future research directions are most promising for advancing the state-of-the-art in chemistry LLMs?
Key findings
- Generalist LLMs underperform on chemistry tasks due to insufficient domain knowledge, leading to errors in reaction prediction, property calculation, and nomenclature.
- Multi-modal LLMs such as 3D-MoLM and ChemVLM show improved performance on 3D structure and image-based tasks, but remain limited by data scarcity and annotation quality.
- SFT-based models like LlaSMol and ChemDFM achieve strong performance on molecular property prediction and reaction generation, with accuracy improvements over zero-shot prompting.
- Tool-augmented LLMs such as ChemCrow and LLaMP demonstrate significant gains in complex tasks like retrosynthesis and reagent prediction by integrating external tools.
- Benchmarks like MassSpecGym and ScholarChemQA reveal that LLMs still struggle with spectral interpretation and domain-specific reasoning, especially under distribution shift.
- The survey identifies a critical gap in open, large-scale, multi-modal datasets—especially for realistic molecular images and reaction schematics—limiting model generalization.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.