Skip to main content
QUICK REVIEW

[논문 리뷰] A Review of Large Language Models and Autonomous Agents in Chemistry

Mayk Caldas Ramos, Christopher J. Collison|arXiv (Cornell University)|2024. 06. 26.
Machine Learning in Materials Science인용 수 12
한 줄 요약

이 리뷰는 대형 언어 모델(LLMs)과 LLM 기반 자율 에이전트가 화학 분야에 미치는 영향을 다루며, 아키텍처, 화학 응용, 데이터 세트, 벤치마크, 도전 과제 및 향후 방향을 포괄합니다.

ABSTRACT

Large language models (LLMs) have emerged as powerful tools in chemistry, significantly impacting molecule design, property prediction, and synthesis optimization. This review highlights LLM capabilities in these domains and their potential to accelerate scientific discovery through automation. We also review LLM-based autonomous agents: LLMs with a broader set of tools to interact with their surrounding environment. These agents perform diverse tasks such as paper scraping, interfacing with automated laboratories, and synthesis planning. As agents are an emerging topic, we extend the scope of our review of agents beyond chemistry and discuss across any scientific domains. This review covers the recent history, current capabilities, and design of LLMs and autonomous agents, addressing specific challenges, opportunities, and future directions in chemistry. Key challenges include data quality and integration, model interpretability, and the need for standard benchmarks, while future directions point towards more sophisticated multi-modal agents and enhanced collaboration between agents and experimental methods. Due to the quick pace of this field, a repository has been built to keep track of the latest studies: https://github.com/ur-whitelab/LLMs-in-science.

연구 동기 및 목표

  • LLM이 화학에서 특성 예측, 역 설계, 합성 계획을 가능하게 하는 방식을 평가한다.
  • 화학 과제에 대한 인코더-전용, 디코더-전용, 인코더-디코더 LLM 아키텍처를 비교한다.
  • LLM 기반 자율 에이전트와 이들의 문헌 조사, 실험, 데이터 자동화에서의 역할을 논의한다.
  • 데이터 품질, 벤치마크, 해석 가능성 및 통합 과제를 식별하고 방향을 제시한다.

제안 방법

  • 트랜스포머의 역사적 맥락을 제공하고 아키텍처를 화학 과제에 매핑한다.
  • 화학 LLM에 관련된 분자 표현, 데이터 세트 및 벤치마크를 검토한다.
  • 특성 예측, 합성 및 다중 모달 작업을 위한 LLM 유형(인코더-전용, 디코더-전용, 인코더-디코더)을 분석한다.
  • 사전 학습, 지도 미세 조정, RLHF, DPO 등 학습 파이프라인 및 정렬 방법을 논의한다.
  • 메모리, 계획, 지각, 도구를 포함한 자율 에이전트 설계와 그 화학 응용을 조사한다.
  • 다중 모달 에이전트 및 에이전트-실험 방법 협력의 향후 방향을 종합한다.
Figure 1: a) The generalized encoder-decoder transformer: The encoder on the left converts an input into a vector, while the decoder on the right predicts the next token in a sequence. b) Encoder-decoder transformers are traditionally used for translation tasks and, in chemistry, for reaction predic
Figure 1: a) The generalized encoder-decoder transformer: The encoder on the left converts an input into a vector, while the decoder on the right predicts the next token in a sequence. b) Encoder-decoder transformers are traditionally used for translation tasks and, in chemistry, for reaction predic

실험 결과

연구 질문

  • RQ1화학에서 특성 예측, 역 설계, 합성에 대한 현재 LLM의 능력은 무엇인가?
  • RQ2서로 다른 트랜스포머 아키텍처가 화학 특화 작업에서 어떻게 수행하는가?
  • RQ3화학 LLM을 가장 잘 지원하는 데이터 품질, 벤치마크 및 분자 표현은 무엇인가?
  • RQ4화학에서 LLM 기반 자율 에이전트의 주된 도전 과제와 기회는 무엇인가?
  • RQ5다중 모달 및 협력 에이전트의 향후 개발이 실험 화학을 어떻게 발전시킬 수 있는가?

주요 결과

  • LLMs enable property prediction, molecule design, and synthesis planning by leveraging chemistry language representations like SMILES and InChI.
  • Encoder-only models (e.g., BERT-based) excel at property prediction and reaction classification, while decoder-only models enable de novo molecule generation, and encoder–decoder hybrids support flexible tasks.
  • Multi-modal and text-to-text approaches (e.g., T5, instruction tuning) broaden task scope and generalization in chemical domains.
  • Autonomous agents equipped with memory, planning, perception, and tool interfaces can conduct literature review, experiment planning, and automated data processing in chemistry.
  • A critical bottleneck is data quality and grounding; existing datasets (e.g., MoleculeNet) have limitations, underscoring a need for high-quality, real-world grounded data and standardized benchmarks.
  • Future directions point to more sophisticated multi-modal agents and tighter coupling between agents and experimental laboratories.
Figure 3: Number of training tokens (on log scale) available from various chemical sources compared with typical LLM training runs. The numbers are drawn from ZINC 176 , PubChem 177 , Touvron et al. 178 , ChEMBL 179 , and Kinney et al. 180
Figure 3: Number of training tokens (on log scale) available from various chemical sources compared with typical LLM training runs. The numbers are drawn from ZINC 176 , PubChem 177 , Touvron et al. 178 , ChEMBL 179 , and Kinney et al. 180

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.