Skip to main content
QUICK REVIEW

[논문 리뷰] Accurate and efficient structure elucidation from routine one-dimensional NMR spectra using multitask machine learning

Frank Hu, Michael S. Chen|arXiv (Cornell University)|2024. 08. 15.
Molecular spectroscopy and chiralityChemistry인용 수 3
한 줄 요약

이 논문은 분자의 화학식이나 구조 브리징 정보 없이도 원시 1D 1H 및 13C NMR 스펙트럼에서 유기 분자의 정확한 종단 간 구조 규명을 가능하게 하는 다중 작업 기계학습 프레임워크를 제시한다. 하위구조에서 구조로의 매핑을 사전 훈련한 트랜스포머 기반 아키텍처와 스펙트럼 인코딩을 위한 CNN을 통합함으로써, 최대 19개의 무거운 원자를 가진 분자에 대해 정상적인 15위 정확도가 69.6%에 도달하며, 검색 공간을 최대 11개 주자 정도로 감소시킨다.

ABSTRACT

Rapid determination of molecular structures can greatly accelerate workflows across many chemical disciplines. However, elucidating structure using only one-dimensional (1D) NMR spectra, the most readily accessible data, remains an extremely challenging problem because of the combinatorial explosion of the number of possible molecules as the number of constituent atoms is increased. Here, we introduce a multitask machine learning framework that predicts the molecular structure (formula and connectivity) of an unknown compound solely based on its 1D 1H and/or 13C NMR spectra. First, we show how a transformer architecture can be constructed to efficiently solve the task, traditionally performed by chemists, of assembling large numbers of molecular fragments into molecular structures. Integrating this capability with a convolutional neural network (CNN), we build an end-to-end model for predicting structure from spectra that is fast and accurate. We demonstrate the effectiveness of this framework on molecules with up to 19 heavy (non-hydrogen) atoms, a size for which there are trillions of possible structures. Without relying on any prior chemical knowledge such as the molecular formula, we show that our approach predicts the exact molecule 69.6% of the time within the first 15 predictions, reducing the search space by up to 11 orders of magnitude.

연구 동기 및 목표

  • 일반적인 1D NMR 스펙트럼에서 분자의 화학식이나 조각 정보에 의존하지 않고도 비지도적이고 정확한 구조 규명 문제를 해결하기 위해.
  • 분자의 크기가 10~19개의 무거운 원자를 넘어서면서 가능한 분자 구조의 조합 폭발 문제를 해결하기 위해.
  • 원시 NMR 스펙트럼에서 직접 분자 결합성과 화학식으로 이어지는 종단 간 딥 러닝 프레임워크를 개발하기 위해.
  • 빠르고 확장 가능하며 접근 가능한 구조 규명을 가능하게 하여 화학 연구, 교육, 산업 워크플로우 전반에 적용 가능하게 하기 위해.
  • 향후 입체화학, 더 큰 분자, 더 넓은 원소 다양성으로의 확장 기반을 마련하기 위해.

제안 방법

  • 957개의 단순 하위구조(≤7개 원자)의 존재 또는 부재를 기반으로 분자 구조를 재구성하는 데 목적이 있는 트랜스포머 모델을 사전 훈련하여 효율적인 구조 조립을 가능하게 한다.
  • 기존의 광범위한 전처리 없이도 원시 1D 1H 및 13C NMR 스펙트럼을 잠재 표현으로 인코딩하기 위해 컨volutional 신경망(CNN)을 사용한다.
  • 사전 훈련된 트랜스포머와 CNN을 통합하여 스펙트럼에서 분자의 하위구조와 전체 분자 구조를 동시에 예측하는 다중 작업 학습 프레임워크를 구성한다.
  • 화학식이 아닌 스펙트럼 데이터만을 입력으로 사용하여 최대 19개의 무거운 원자를 가진 분자의 시뮬레이션된 NMR 스펙트럼을 기반으로 종단 간으로 훈련한다.
  • 후보 구조를 생성하고 순위를 매기기 위해 빔 서치 전략을 활용하며, 상위 예측 결과를 정확도 평가에 사용한다.
  • 10~19개의 무거운 원자를 가진 분자에 대한 벤치마크에서 상위 1위 및 상위 15위 정확도를 측정하여 모델을 평가한다.
Figure 1: Overview of the full multitask structure elucidation workflow (top) and the substructure-to-structure workflow (bottom). Weights from a transformer pretrained on the substructure-to-structure task are used to initialize the multitask model. Specific details regarding the transformer model
Figure 1: Overview of the full multitask structure elucidation workflow (top) and the substructure-to-structure workflow (bottom). Weights from a transformer pretrained on the substructure-to-structure task are used to initialize the multitask model. Specific details regarding the transformer model

실험 결과

연구 질문

  • RQ1원시 1D NMR 스펙트럼에서 분자의 화학식이나 구조 브리징 정보 없이도 다중 작업 딥 러닝 모델이 분자 구조를 정확하게 예측할 수 있는가?
  • RQ2분자 크기가 증가함에 따라 가능한 분자 수가 조합적으로 증가할 때 모델의 성능은 어떻게 변화하는가?
  • RQ3트랜스포머 기반 아키텍처는 하위구조 존재/부재 신호로부터 분자 결합성을 효과적으로 재구성할 수 있는가?
  • RQ4CNN를 통한 스펙트럼 인코딩과 트랜스포머를 통한 구조 생성의 통합이 이전 방법에 비해 예측 정확도를 얼마나 향상시키는가?
  • RQ5복잡한 스펙트럼 해석에 비해 다양한 분자 스케일드와 구조를 통해 일반화할 수 있는가, 높은 정확도를 유지하는가?

주요 결과

  • 모델은 원시 1H 및 13C NMR 스펙트럼만을 사용하여 최대 19개의 무거운 원자를 가진 분자에 대해 정상적인 15위 정확도가 69.6%에 도달한다.
  • 이 프레임워크는 효과적인 검색 공간을 최대 11개 주자 정도로 감소시켜 트리리온의 가능한 구조 탐색을 효율적으로 가능하게 한다.
  • 모델은 분자 크기에 관계없이 높은 성능를 유지하며, 가능한 분자 수가 5개 주자 증가할 때 정확도가 25.5% 감소하는 데 그친다.
  • 사전 훈련된 트랜스포머만으로도 하위구조 입력에서의 구조 재구성에 대해 93.2%의 상위 15위 정확도를 달성하여 분자 조립에 있어 뛰어난 강건성을 입증한다.
  • 표준 CPU(AMD Ryzen 7 3700X)에서 전체 구조 예측이 3초 이내로 완료되어 실생활 적용에 매우 접근 가능하고 실용적이다.
  • 모델은 일반화 가능하며, 훈련 데이터에 입체 중심 및 이중 결합 구조 정보를 통합함으로써 입체화학 예측으로의 확장이 가능하다.
Figure 2: (Left) Transformer and the best multitask model test accuracy as a function of the problem size. The problem size is determined by extrapolating an exponential fit to the number of molecules in GDB-9 33 , GDB-11 34 , GDB-13 35 , and GDB-17, and the plot begins with the number of possible s
Figure 2: (Left) Transformer and the best multitask model test accuracy as a function of the problem size. The problem size is determined by extrapolating an exponential fit to the number of molecules in GDB-9 33 , GDB-11 34 , GDB-13 35 , and GDB-17, and the plot begins with the number of possible s

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.