[Paper Review] A Review of Large Language Models and Autonomous Agents in Chemistry
This review surveys how large language models (LLMs) and LLM-based autonomous agents are shaping chemistry, covering architectures, chemistry applications, datasets, benchmarks, challenges, and future directions.
Large language models (LLMs) have emerged as powerful tools in chemistry, significantly impacting molecule design, property prediction, and synthesis optimization. This review highlights LLM capabilities in these domains and their potential to accelerate scientific discovery through automation. We also review LLM-based autonomous agents: LLMs with a broader set of tools to interact with their surrounding environment. These agents perform diverse tasks such as paper scraping, interfacing with automated laboratories, and synthesis planning. As agents are an emerging topic, we extend the scope of our review of agents beyond chemistry and discuss across any scientific domains. This review covers the recent history, current capabilities, and design of LLMs and autonomous agents, addressing specific challenges, opportunities, and future directions in chemistry. Key challenges include data quality and integration, model interpretability, and the need for standard benchmarks, while future directions point towards more sophisticated multi-modal agents and enhanced collaboration between agents and experimental methods. Due to the quick pace of this field, a repository has been built to keep track of the latest studies: https://github.com/ur-whitelab/LLMs-in-science.
Motivation & Objective
- Assess how LLMs enable property prediction, inverse design, and synthesis planning in chemistry.
- Compare encoder-only, decoder-only, and encoder-decoder LLM architectures for chemistry tasks.
- Discuss LLM-based autonomous agents and their roles in literature review, experiments, and data automation.
- Identify data quality, benchmarks, interpretability, and integration challenges and propose directions.
Proposed method
- Provide historical context of transformers and map architectures to chemistry tasks.
- Review molecular representations, datasets, and benchmarks relevant to chemistry LLMs.
- Analyze LLM types (encoder-only, decoder-only, encoder-decoder) for property prediction, synthesis, and multi-modal tasks.
- Discuss training pipelines and alignment methods (pretraining, supervised fine-tuning, RLHF, DPO).
- Survey autonomous agents design (memory, planning, perception, tools) and their chemistry applications.
- Synthesize future directions for multi-modal agents and agent–experimental-method collaboration.

Experimental results
Research questions
- RQ1What are the current capabilities of LLMs in property prediction, inverse design, and synthesis for chemistry?
- RQ2How do different transformer architectures perform on chemistry-specific tasks?
- RQ3What data quality, benchmarks, and molecular representations best support chemistry LLMs?
- RQ4What are the main challenges and opportunities for LLM-based autonomous agents in chemistry?
- RQ5What future developments in multi-modal and collaborative agents could advance experimental chemistry?
Key findings
- LLMs enable property prediction, molecule design, and synthesis planning by leveraging chemistry language representations like SMILES and InChI.
- Encoder-only models (e.g., BERT-based) excel at property prediction and reaction classification, while decoder-only models enable de novo molecule generation, and encoder–decoder hybrids support flexible tasks.
- Multi-modal and text-to-text approaches (e.g., T5, instruction tuning) broaden task scope and generalization in chemical domains.
- Autonomous agents equipped with memory, planning, perception, and tool interfaces can conduct literature review, experiment planning, and automated data processing in chemistry.
- A critical bottleneck is data quality and grounding; existing datasets (e.g., MoleculeNet) have limitations, underscoring a need for high-quality, real-world grounded data and standardized benchmarks.
- Future directions point to more sophisticated multi-modal agents and tighter coupling between agents and experimental laboratories.

Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.