[논문 리뷰] TopP-S: Persistent homology based multi-task deep neural networks for simultaneous predictions of partition coefficient and aqueous solubility
이 논문은 소분자의 수용성과 분배계수(logP)를 동시에 예측하기 위해 요소별 지속 호몰로지(ESPH)를 사용하는 다중 작업 딥러닝 프레임워크인 TopP-S를 제안한다. 분자의 위상적 구조를 스케일러블하고 다중 척도의 위상적 불변량으로 인코딩함으로써, 상관된 물리화학적 성질을 공동으로 학습함으로써 벤치마크 데이터셋에서 최신 기술을 능가하는 높은 정확도의 예측을 가능하게 한다.
Aqueous solubility and partition coefficient are important physical properties of small molecules. Accurate theoretical prediction of aqueous solubility and partition coefficient plays an important role in drug design and discovery. The prediction accuracy depends crucially on molecular descriptors which are typically derived from theoretical understanding of the chemistry and physics of small molecules. The present work introduces an algebraic topology based method, called element specific persistent homology (ESPH), as a new representation of small molecules that is entirely different from conventional chemical and/or physical representations. ESPH describes molecular properties in terms of multiscale and multicomponent topological invariants. Such topological representation is systematical, comprehensive, and scalable with respect to molecular size and composition variations. However, it cannot be literally translated into a physical interpretation. Fortunately, it is readily suitable for machine learning methods, rendering topological learning algorithms. Due to the inherent correlation between solubility and partition coefficient, a uniform ESPH representation is developed for both properties, which facilitates multi-task deep neural networks for their simultaneous predictions. This strategy leads to more accurate prediction of relatively small data sets. A total of six data sets is considered in the present work to validate the proposed topological and multi-task deep learning approaches. It is demonstrate that the proposed approaches achieve some of the most accurate predictions of aqueous solubility and partition coefficient. Our software is available online at {\url{http://weilab.math.msu.edu/TopP-S/}}
연구 동기 및 목표
- 다양한 척도와 다성분 구조적 특징을 포괄하는 대수적 위상기반의 새로운 분자 표현 방식을 개발한다.
- 약물 설계에서 핵심적인 성질인 logP와 수용성을 정확하게 예측하는 데 도전하며, 자료 효율적이고 위상 정보를 반영한 기계학습 기법을 적용한다.
- logP와 수용성 간의 내재된 상관관계를 활용하여 다중 작업 학습을 통해 두 성질을 함께 모델링한다.
- 지속 호몰로지에서 유도된 위상 기반 기술자가 전통적인 분자 기술자보다 예측 정확도에서 뛰어나다는 것을 입증한다.
제안 방법
- 분자의 기하학적 구조에서 원소별 지속 호몰로지(ESPH)를 활용하여 원자 종류와 결합 패턴을 인코딩한 다중 척도의 다성분 위상 불변량을 생성한다.
- ESPH 출력의 압축되고 해석 가능한 요약으로서 원소별 위상 기술자(ESTD)를 구축하여 화학적 및 위상적 정보를 유지한다.
- ESTD를 딥뉴럴넷과 통합하여 logP와 수용성을 동시에 예측하는 다중 작업 학습 프레임워크를 구현한다.
- 예측 성능 평가를 위해 다양한 데이터셋에서 기울기 부스팅, 랜덤 포레스트 등의 앙상블 방법과 딥뉴럴넷을 적용한다.
- 모델의 일반화 능력과 강건성을 평가하기 위해 10겹 교차검증과 한 개씩 제거하는 검증(leave-one-out)을 실시한다.
- 공유된 표현을 통해 관련된 성질에서 유도된 인덕티브 바이어스를 활용하여, 소규모 또는 제한된 데이터셋에서의 성능 향상을 도모한다.
실험 결과
연구 질문
- RQ1원소별 지속 호몰로지(ESPH)는 성질 예측에 필수적인 화학적·물리적 특징을 체계적이고 확장 가능하며 정보적인 방식으로 표현할 수 있는가?
- RQ2logP와 수용성을 공동으로 다중 작업 학습하는 방식이 단일 작업 모델 대비 예측 정확도를 어떻게 향상시키는가?
- RQ3위상 기반 기술자(ESTD)가 전통적인 2차원 분자 기술자보다 logP와 수용성을 예측하는 데 얼마나 더 뛰어난가?
- RQ4직접적인 물리적 해석이 없는 위상 기반 표현이, 벤치마크 데이터셋에서 최첨단 성능을 달성할 수 있는가?
주요 결과
- ESPH 기반의 위상 표현을 ESTD로 인코딩한 방식은 여섯 개의 벤치마크 데이터셋에서 logP와 수용성 예측에 있어 가장 높은 정확도를 달성한다.
- 공유된 ESTD 표현을 기반으로 훈련된 다중 작업 딥뉴럴넷(MT-DNN)은 특히 소규모 데이터셋에서 단일 작업 모델보다 유의미하게 뛰어난 성능을 보이며, 공유된 인덕티브 바이어스 덕분에 일반화 능력이 향상된다.
- FDA 및 Star 데이터셋에서 logP 예측 성능이 최첨단 수준에 도달했으며, 보고된 RMSE 값은 0.5 log 단위 이하였다.
- 기존의 2차원 기술자와 함께 ESTD를 포함시킴으로써 추가적인 예측 정확도 향상이 이루어졌으며, 이는 상호보완적인 정보의 유용성을 입증한다.
- 다양한 데이터 분포에서의 강건성을 입증하였으며, 교차검증 및 테스트 세트 평가에서 일관된 성능 향상이 관찰되었다.
- http://weilab.math.msu.edu/TopP-S/ 에서 제공하는 온라인 서버를 통해 모델에 접근 가능하여 재현성과 약물 발견 파이프라인 내 실용적 적용이 가능하다.
더 나은 연구,지금 바로 시작하세요
논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.
카드 등록 없음 · 무료 플랜 제공
이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.