[논문 리뷰] From Words to Molecules: A Survey of Large Language Models in Chemistry
이 연구 조사는 대형 언어 모델(LLMs)이 화학에 어떻게 적용되는지 분류하고, 분자 표현, 토큰화, 사전학습 목표 및 응용 패러다임을 자세히 설명한다. 또한 향후 연구 방향을 제시한다.
In recent years, Large Language Models (LLMs) have achieved significant success in natural language processing (NLP) and various interdisciplinary areas. However, applying LLMs to chemistry is a complex task that requires specialized domain knowledge. This paper provides a thorough exploration of the nuanced methodologies employed in integrating LLMs into the field of chemistry, delving into the complexities and innovations at this interdisciplinary juncture. Specifically, our analysis begins with examining how molecular information is fed into LLMs through various representation and tokenization methods. We then categorize chemical LLMs into three distinct groups based on the domain and modality of their input data, and discuss approaches for integrating these inputs for LLMs. Furthermore, this paper delves into the pretraining objectives with adaptations to chemical LLMs. After that, we explore the diverse applications of LLMs in chemistry, including novel paradigms for their application in chemistry tasks. Finally, we identify promising research directions, including further integration with chemical knowledge, advancements in continual learning, and improvements in model interpretability, paving the way for groundbreaking developments in the field.
연구 동기 및 목표
- 분자 정보가 LLM에 대해 어떻게 토큰화되고 표현되는지 체계적으로 검토한다.
- 입력 도메인과 모달리티를 기준으로 한 화학 LLM의 분류 체계를 제시한다(단일 도메인, 다중 도메인, 다중 모달).
- 화학 데이터에 대한 사전학습 목표 및 적응 전략을 분석한다.
- LLMs가 가능하게 하는 다양한 화학 응용을 탐구하고 열려있는 연구 방향을 식별한다.
제안 방법
- 분자 표현(지문, SMILES/SELFIES, InChI, 그래프 기반)과 토큰화 수준(character-, atom-, motif-level)을 분류한다.
- 사전학습 데이터 도메인(단일 도메인, 다중 도메인, 다중 모달) 및 통합 전략의 분류 체계를 제시한다.
- 화학 LLM을 위한 핵심 사전학습 목표 세 가지를 검토한다: Masked Language Modeling (MLM), Molecule Property Prediction (MPP), 및 Autoregressive Token Generation (ATG)과 화학 특화 작업.
- Cross-modal contrastive learning (XMC) 및 모달리티 간 정렬과 같은 교차 모달 목표를 논의한다.
- 방법의 포괄적 표에서 대표적인 아키텍처, 데이터셋 및 학습 접근법을 요약한다.
- 지속적인 학습 및 해석 가능성을 포함한 응용 및 향후 방향을 개략한다.
실험 결과
연구 질문
- RQ1화학에서 분자 시퀀스가 LLM에 대해 어떻게 토큰화되고 표현되는가?
- RQ2입력 도메인과 모달리티에 따라 화학 LLM을 가장 잘 포착하는 분류 체계는 무엇인가?
- RQ3무슨 사전학습 목표가 사용되며 화학 데이터에 어떻게 적용/적합화되는가?
- RQ4화학 LLM이 가능하게 하는 주요 응용 및 패러다임은 무엇인가?
- RQ5화학 지식의 통합, 지속 학습 및 해석 가능성을 향상시킬 수 있는 향후 방향은 무엇인가?
주요 결과
- 분자 표현은 지문, SMILES/SELFIES, InChI, 그리고 다양한 입자(그래프 기반) 형태를 포함한다.
- 토큰화 체계는 문자 수준, 원자 수준, 모티프 수준으로 확장되며 데이터 기반 및 화학 주도 방법을 포함한다.
- 화학 LLM은 입력 데이터와 모달리티에 따라 단일 도메인, 다중 도메인, 다중 모달의 분류로 조직된다.
- MLM, MPP, and ATG가 핵심 사전학습 목표이며, MPP는 강한 표현 학습 신호를 제공하고 ATG는 작업 정렬을 가능하게 한다.
- Cross-modal 학습 및 표현 정렬은 텍스트, 그래프, 지문, 이미지를 융합하는 데 사용되며, 도메인 특유의 뉘앙스는 여전히 도전 과제로 남아 있다.
- 응용 분야에는 챗봇, 컨텍스트 내 학습, 및 특성 예측, 반응 예측, 분자 생성 등 다운스트림 작업을 위한 표현 학습이 포함된다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.