Skip to main content
QUICK REVIEW

[논문 리뷰] IMG2SMI: Translating Molecular Structure Images to Simplified Molecular-input Line-entry System

Daniel Campos, Heng Ji|arXiv (Cornell University)|2021. 09. 03.
Biomedical Text Mining and Ontologies참고 문헌 52인용 수 9
한 줄 요약

IMG2SMI는 2D 분자 구조 이미지를 SMILES 문자열로 변환하는 딥러닝 모델을 제안한다. 이 모델은 특징 추출을 위해 ResNet-101 백본을 사용하고, 시퀀스 생성을 위해 Transformer 기반의 인코더-디코더 아키텍처를 활용한다. OSRA 대비 분자 유사도 예측에서 163% 향상되었으며, MACCS Tanimoto 유사도는 0.9475를 기록했고, 벤치마킹을 위해 새로운 8100만 개 분자로 구성된 데이터셋(MOLCAP)을 제공한다.

ABSTRACT

Like many scientific fields, new chemistry literature has grown at a staggering pace, with thousands of papers released every month. A large portion of chemistry literature focuses on new molecules and reactions between molecules. Most vital information is conveyed through 2-D images of molecules, representing the underlying molecules or reactions described. In order to ensure reproducible and machine-readable molecule representations, text-based molecule descriptors like SMILES and SELFIES were created. These text-based molecule representations provide molecule generation but are unfortunately rarely present in published literature. In the absence of molecule descriptors, the generation of molecule descriptors from the 2-D images present in the literature is necessary to understand chemistry literature at scale. Successful methods such as Optical Structure Recognition Application (OSRA), and ChemSchematicResolver are able to extract the locations of molecules structures in chemistry papers and infer molecular descriptions and reactions. While effective, existing systems expect chemists to correct outputs, making them unsuitable for unsupervised large-scale data mining. Leveraging the task formulation of image captioning introduced by DECIMER, we introduce IMG2SMI, a model which leverages Deep Residual Networks for image feature extraction and an encoder-decoder Transformer layers for molecule description generation. Unlike previous Neural Network-based systems, IMG2SMI builds around the task of molecule description generation, which enables IMG2SMI to outperform OSRA-based systems by 163% in molecule similarity prediction as measured by the molecular MACCS Fingerprint Tanimoto Similarity. Additionally, to facilitate further research on this task, we release a new molecule prediction dataset. including 81 million molecules for molecule description generation

연구 동기 및 목표

  • 화학 문헌은 주로 시각적 자료이며 SMILES 문자열을 포함하지 않기 때문에, 이를 기계로 읽을 수 있는 분자 표현으로 추출하는 데 도전한다.
  • 수동 보정이 필요하고 대규모 비지도 탐색에 부적합한 수작업 규칙 기반 시스템(예: OSRA, ChemSchematicResolver)의 한계를 극복한다.
  • 직접 분자 구조 이미지에서 정확한 SMILES 문자열을 생성할 수 있는 종단간 신경망 모델을 개발한다.
  • 향후 화학문헌 분석 및 분자 표현 학습 연구를 지원하기 위해 대규모 공개 데이터셋(MOLCAP)을 제공한다.

제안 방법

  • 2D 분자 구조 이미지에서 시각적 특징을 추출하기 위해 ResNet-101 백본을 활용한다.
  • 추출된 이미지 특징에서 SMILES 문자열을 생성하기 위해 Transformer 기반의 인코더-디코더 아키텍처를 사용한다.
  • 이미지 캡션 생성 파рад림을 활용하여 분자 기술 생성 작업에 적합하게 변형한다.
  • 이미지 스타일, 회전, 레이아웃를 다양하게 조절하기 위해 RDKIT를 사용해 SMILES 문자열에서 생성한 8100만 개 분자의 합성 데이터셋을 훈련에 활용한다.
  • 분포 이탈에 대한 강건성을 향상시키기 위해 자르기, 회전, 노이즈 주입 등의 데이터 증강 기법을 적용한다.
  • 기존의 문서 세그멘테이션 도구와 모델을 조합하여 과학 논문에서 종단간 분자 추출을 가능하게 한다.

실험 결과

연구 질문

  • RQ1수작업 규칙에 의존하지 않고 합성 분자 이미지에서 훈련된 딥러닝 모델이 정확한 SMILES 문자열을 생성할 수 있는가?
  • RQ2분자 기술 생성을 위한 시각-언어 모델의 성능은 기존 수작업 규칙 기반 시스템(예: OSRA)과 비교해 분자 유사도 측면에서 어떻게 다른가?
  • RQ3데이터 증강 및 모델 아키텍처 선택이 다양한 분자 구조와 이미지 변형에 대한 일반화 능력 향상에 얼마나 기여하는가?
  • RQ4대규모로 공개된 분자 이미지 및 SMILES 쌍 데이터셋은 화학 AI에서의 사전 훈련 및 전이 학습 향상에 기여할 수 있는가?

주요 결과

  • MACCS Tanimoto 유사도 지표를 사용해 OSRA 대비 163% 높은 분자 유사도 예측 성능을 기록했으며, 점수는 0.9475를 달성했다.
  • ROUGE-L 점수는 0.63을 기록해 토큰의 강한 재현성을 보였지만, 추가 또는 잘못된 토큰 포함로 인해 정밀도는 낮았다.
  • 정확한 일치 정확도는 7.24%에 불과해 의미적 유사도와 문법적 정확성 사이에 뚜렷한 격차가 있음을 시사한다.
  • 지문 기반 지표(MACCS, Morgan, Path)는 일관된 방향성 추세를 보였지만, 하위구조 수가 적어 민감도가 낮은 MACCS를 제외한 다른 지표와 크기에서 차이를 보였다.
  • 단백질 길이가 짧은 분자(<40 토큰)에서는 기존 시스템인 OSRA보다 성능이 열 劣하므로 하이브리드 접근이 필요하다는 점을 시사한다.
  • 8100만 개 분자-이미지 쌍으로 구성된 MOLCAP 데이터셋의 공개는 향후 화학문헌 AI 연구를 위한 대규모 벤치마크를 제공한다.

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.