Skip to main content
QUICK REVIEW

[논문 리뷰] Embracing assay heterogeneity with neural processes for markedly improved bioactivity predictions

Lucian Chan, Marcel L. Verdonk|arXiv (Cornell University)|2023. 08. 17.
Computational Drug Discovery MethodsComputer Science인용 수 3
한 줄 요약

이 논문은 생물활성 데이터의 시험 이질성에 대해 신경 과정을 사용하는 메타학습 프레임워크인 MetaBind을 제안하며, 다양한 단백질 타겟과 시험 유형 간에 매우 정확하고 유연한 친화도 예측을 가능하게 한다. 희소하고 이질적인 데이터에서 학습하고 태스크별 지지 세트를 활용함으로써 MetaBind는 기존 모델보다 훨씬 낮은 오차율을 기록하며 최신 기술 수준의 성능을 달성한다.

ABSTRACT

Predicting the bioactivity of a ligand is one of the hardest and most important challenges in computer-aided drug discovery. Despite years of data collection and curation efforts by research organizations worldwide, bioactivity data remains sparse and heterogeneous, thus hampering efforts to build predictive models that are accurate, transferable and robust. The intrinsic variability of the experimental data is further compounded by data aggregation practices that neglect heterogeneity to overcome sparsity. Here we discuss the limitations of these practices and present a hierarchical meta-learning framework that exploits the information synergy across disparate assays by successfully accounting for assay heterogeneity. We show that the model achieves a drastic improvement in affinity prediction across diverse protein targets and assay types compared to conventional baselines. It can quickly adapt to new target contexts using very few observations, thus enabling large-scale virtual screening in early-phase drug discovery.

연구 동기 및 목표

  • 약물 발굴에서 희소하고 이질적인 생물활성 데이터의 과제를 해결함으로써 기존 예측 모델의 성능과 이식 가능성에 제한을 둔다.
  • 실험 조건과 시험 유형의 차이에도 불구하고 서로 다른 시험 간의 정보 시너지를 효과적으로 활용할 수 있는 모델을 개발한다.
  • 단지 몇 개의 관측된 활성 측정값만으로도 새로운 타겟과 시험에 빠르게 적응할 수 있도록 하여 대규모 가상 스크리닝을 지원한다.
  • 시험이나 변수를 노이즈로 간주하는 대신 시험 수준의 변동성을 명시적으로 모델링하여 리간드 친화도 예측의 정확성과 강건성을 향상시킨다.

제안 방법

  • 모델은 각 시험이 별개의 학습 태스크로 간주되며, K개의 공분산 클러스터로 그룹화되어 공통적인 구조적 패턴을 모델링하는 계층적 메타학습 프레임워크를 사용한다.
  • 신경 과정 아키텍처를 사용하여, 관측된 리간드-단백질 쌍과 그들의 측정된 생물활성에 기반한 소규모 지지 세트를 조건으로 하여 각 시험에 대한 함수 분포를 학습한다.
  • 리간드 표현을 위해 그래프 컬러지피션 네트워크를, 단백질 임베딩을 위해 컨볼루션 네트워크를 사용하며, 태스크 간 정보를 통합하기 위해 교차 공분산 연산자를 사용한다.
  • 학습 중에 Kullback-Leibler 발산 손실을 적용하여, 각 시험의 전체 공분산 구조를 관측의 무작위 부분집합으로부터 재구성하도록 모델을 유도한다.
  • 프레임워크는 ChEMBL 30 데이터에서 엔드 투 엔드로 훈련되며, Ki, Kd, IC50, EC50 종료 조건을 갖는 고신뢰도, 인간 타겟 기반의 결합 시험에 대해 필터링된다.
  • 모델은 태스크 수준의 지표(T-RMSE, T-MAE)와 쌍화된 시험 분할을 사용하여, 동일한 타겟을 갖는 이질적인 시험 간의 구분 능력을 평가한다.
Figure 1: Illustration of the MetaBind approach. The model constructs an assay-specific SAR using a local support set of per-assay observations to predict affinities for unobserved protein-ligand pairs. In meta-learning terms, each assay is thus treated as a task, and assays are clustered into a sma
Figure 1: Illustration of the MetaBind approach. The model constructs an assay-specific SAR using a local support set of per-assay observations to predict affinities for unobserved protein-ligand pairs. In meta-learning terms, each assay is thus treated as a task, and assays are clustered into a sma

실험 결과

연구 질문

  • RQ1메타학습 모델은 생물활성 예측을 향상시키기 위해 시험 이질성을 효과적으로 캡처하고 활용할 수 있는가?
  • RQ2희소한 레이블 데이터로 새로운, 알려지지 않은 타겟과 시험에 얼마나 잘 일반화되는가?
  • RQ3생물학적으로 유사하지만 실험적으로 이질적인 시험 간의 차이를 모델이 어느 정도 정확하게 식별할 수 있는가?
  • RQ4시험 수준의 변동성을 고려함으로써 기존 모델에 비해 예측 정확도에 상당한 향상이 이루어지는가?

주요 결과

  • MetaBind는 테스트 세트에서 태스크 수준의 평균제곱근오차(T-RMSE)를 0.48 pK 단위로 기록하여 기존의 일반 모델보다 극적으로 향상된 성능을 보였다.
  • 모델는 강력한 제로샷 일반화 능력을 보였으며, 178개의 알려지지 않은 타겟 단백질에서 T-MAE가 0.36 pK 단위로, 훈련 및 테스트 세트 간의 중앙값 리간드 유사도가 0.32였다.
  • 쌍화된 시험 분할에서 MetaBind는 동일한 단백질을 타겟으로 하지만 서로 다른 구조-활성 관계를 보이는 시험 간의 차이를 성공적으로 식별하고 모델링하여, 이질성 처리 능력이 뛰어나다는 것을 보여주었다.
  • 생화학적 및 세포 기반 시험을 포함한 다양한 시험 유형에서 표준 딥러닝 및 가우시안 프로세스 기반 모델보다 뛰어난 성능을 보였다.
  • Kullback-Leibler 손실 항목을 사용함으로써 부분 관측 세트에서 효과적인 공분산 구조 재구성 능력이 향상되어 일반화 및 불확실성 캘리브레이션 능력이 향상되었다.
Figure 2: Observed and predicted heterogeneities in structure-activity relationships. (a-c) Correlation of the bioactivity read-out for congeneric series of three different assay pairs. The pairs (A,B) and (C,D) share the same protein variants, but differ in assay type and conditions. The assays of
Figure 2: Observed and predicted heterogeneities in structure-activity relationships. (a-c) Correlation of the bioactivity read-out for congeneric series of three different assay pairs. The pairs (A,B) and (C,D) share the same protein variants, but differ in assay type and conditions. The assays of

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.