Skip to main content
QUICK REVIEW

[논문 리뷰] GIT-Mol: A Multi-modal Large Language Model for Molecular Science with Graph, Image, and Text

Pengfei Liu, Yiming Ren|arXiv (Cornell University)|2023. 08. 14.
Computational Drug Discovery Methods참고 문헌 47인용 수 8
한 줄 요약

GIT-Mol은 그래프, 이미지 및 텍스트를 융합한 700M 멀티모달 LLM으로, GIT-Former 모달리티 믹서와 Xmodal 프리-트레이닝 전략을 통해 분자 캡션화, 텍스트 기반 분자 생성, 이미지 인식 및 특성 예측을 향상시킵니다.

ABSTRACT

Large language models have made significant strides in natural language processing, enabling innovative applications in molecular science by processing textual representations of molecules. However, most existing language models cannot capture the rich information with complex molecular structures or images. In this paper, we introduce GIT-Mol, a multi-modal large language model that integrates the Graph, Image, and Text information. To facilitate the integration of multi-modal molecular data, we propose GIT-Former, a novel architecture that is capable of aligning all modalities into a unified latent space. We achieve a 5%-10% accuracy increase in properties prediction and a 20.2% boost in molecule generation validity compared to the baselines. With the any-to-language molecular translation strategy, our model has the potential to perform more downstream tasks, such as compound name recognition and chemical reaction prediction.

연구 동기 및 목표

  • 텍스트 기반 LLM이 분자 그래프와 이미지를 완전히 활용하는 데 필요한 한계를 동기부여하고 해결한다.
  • 그래프, 이미지, 텍스트 모달리티를 하나의 잠재 공간으로 통합하기 위해 GIT-Mol을 개발한다.
  • 모달리티 간 융합을 위한 교차 주의 기반의 GIT-Former를 제안하고 아무것도에서 언어로의 번역이 가능하도록 한다.
  • 분자 캡션화, 데 노보 생성, 이미지 인식 및 특성 예측에서의 개선을 입증한다.
  • 각 모달리티와 학습 전략의 기여를 검증하기 위한 차등 실험 및 분석을 제공한다.

제안 방법

  • 그래프, 이미지, 텍스트를 하나의 잠재 공간으로 매핑하는 교차 주의 기반 모달리티 믹서인 GIT-Former를 도입한다.
  • 텍스트에 대해서는 MolT5, 이미지는 Swin Transformer, 그래프는 GIN의 모달리티별 인코더를 사용하고, 생성 작업에는 MolT5 디코더를 활용한다.
  • 모달리티를 정렬하기 위해 Xmodal-Text Matching (XTM) 및 Xmodal-Text Contrastive Learning (XTC)으로 사전 학습한다.
  • 모달리티 번역 작업을 위한 미세조정 중 아무것도에서 언어로의 프롬프트를 적용한다.
  • MoleculeNet-property 작업에서 미세조정하고 언어 기반 출력에 대해 프롬프트 튜닝을 사용한다.
Figure 1: An overview of GIT-Mol . (a) Internal Information , including sequence and graph structure representations, emphasizes inherent chemical properties and simple topology; (b) External Information , e.g., images and text descriptions, provide richer details and help the human understanding; (
Figure 1: An overview of GIT-Mol . (a) Internal Information , including sequence and graph structure representations, emphasizes inherent chemical properties and simple topology; (b) External Information , e.g., images and text descriptions, provide richer details and help the human understanding; (

실험 결과

연구 질문

  • RQ1GIT-Former가 분자 작업을 위한 그래프, 이미지, 텍스트 모달리티를 공유 잠재 공간으로 효과적으로 정렬할 수 있는가?
  • RQ2다중모달 입력이 단일 모달리티에 비해 분자 캡션화, 이미지 기반 인식 및 SMILES 생성에 더 나은가?
  • RQ3XTM과 XTC 학습 전략이 교차 모달 정렬 및 다운스트림 성능에 어떤 영향을 주는가?
  • RQ4프롬프트 학습이 아무것도에서 언어로의 모달리티 번역 및 특성 예측에 어떤 영향을 미치는가?
  • RQ5GIT-Mol의 분자 특성 예측 정확도 및 분자 생성 타당성에 어떤 이득이 있는가?

주요 결과

  • GIT-Mol은 단일 모달리티 기반 기준선보다 모든 지표에서 더 높은 캡션화 성능을 달성한다.
  • 그래프 기반 변형은 일반적으로 SMILES보다 캡션화 지표에서 우수하며, 다중 모달이 두 가지를 다 능가한다.
  • 차폐실험에서 다중 모달리티가 단일 모달리티 대비 10–15%의 향상을 보인다.
  • 데 노보 생성에서 GIT-Mol-캡션+MolT5가 더 높은 타당도(0.928)와 경쟁력 있는 유사도 지표를 보인다.
  • 교차 모달 프리트레이닝(XTM 우선, 그다음 XTC)과 프롬프트 학습이 결과에 큰 영향을 미친다.
  • 다수의 모달리티 교차 모듈에서 베이스라인 대비 GIT-Mol이 분자 생성 및 특성 예측 작업에서 여러 지표에서 우수하다.
Figure 2: Study case of Molecule Caption . The GIT-Mol model exhibits precise chemical characterization, aligning closely with ground truth information.
Figure 2: Study case of Molecule Caption . The GIT-Mol model exhibits precise chemical characterization, aligning closely with ground truth information.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.