Skip to main content
QUICK REVIEW

[논문 리뷰] Deep Confident Steps to New Pockets: Strategies for Docking Generalization

Gabriele Corso, Arthur Deng|PubMed|2024. 02. 28.
Protein Structure and Dynamics참고 문헌 30인용 수 19
한 줄 요약

이 논문은 Blind docking 일반화용 새로운 벤치마크 DockGen을 소개하고, ML 도킹 모델이 보지 못한 포켓에 일반화가 잘 되지 않는 문제를 보여주며, 일반화 향상 및 최첨단 성능 달성을 위해 스케일링과 Confidence Bootstrapping을 제시합니다.

ABSTRACT

Accurate blind docking has the potential to lead to new biological breakthroughs, but for this promise to be realized, docking methods must generalize well across the proteome. Existing benchmarks, however, fail to rigorously assess generalizability. Therefore, we develop DockGen, a new benchmark based on the ligand-binding domains of proteins, and we show that existing machine learning-based docking models have very weak generalization abilities. We carefully analyze the scaling laws of ML-based docking and show that, by scaling data and model size, as well as integrating synthetic data strategies, we are able to significantly increase the generalization capacity and set new state-of-the-art performance across benchmarks. Further, we propose Confidence Bootstrapping, a new training paradigm that solely relies on the interaction between diffusion and confidence models and exploits the multi-resolution generation process of diffusion models. We demonstrate that Confidence Bootstrapping significantly improves the ability of ML-based docking methods to dock to unseen protein classes, edging closer to accurate and generalizable blind docking methods.

연구 동기 및 목표

  • 다양한 단백질 포켓 전반에서 Blind docking의 일반화 평가 필요성에 대한 동기 부여.
  • DockGen을 학습 포켓을 넘어 교차 도메인 일반화를 평가하는 도메인 수준 벤치마크로 제안.
  • ML 기반 도킹의 일반화에 대한 데이터/모델 증가 효과를 이해하기 위한 스케일링 법칙 분석.
  • 포켓 다양성 확장을 위한 합성 데이터 확장을 도입하고 그 영향 연구.
  • 확정 모델과 확산 모델 간 피드백으로 일반화를 향상시키는 Confidence Bootstrapping 도입

제안 방법

  • ECOD에 따른 단백질 도메인 클러스터링 및 Binding MOAD의 141/189 복합체에 대해 필터링된 검증/테스트 세트를 구성하여 DockGen 벤치마크를 개발.
  • DockGen에서 기초 ML 및 탐색 기반 도킹 방법을 평가하여 일반화 격차를 정량화.
  • 일반화를 연구하기 위해 데이터 및 모델 크기를 확장; 학습을 보강하기 위한 van der Mer 영감을 받은 합성 사이드체인 리간드를 도입.
  • DockGen에서 새로운 SOTA를 달성하기 위해 더 큰 모델과 보강된 데이터를 결합한 DiffDock-L 제안(및 다른 벤치마크 비교).
  • Confidence Bootstrapping 도입: 확산 모델이 포즈를 생성하고, 확신 모델이 이를 평가한 뒤 피드백으로 초기 확산 단계를 업데이트하는 자체 학습 스킴.
  • 반복적 최적화를 통해 확산 점수는 롤아웃과 확신 기반 재가중화를 통해 반복적으로 업데이트됩니다.
Figure 1: Visual representation of the Confidence Bootstrapping training scheme. The dashed lines represent the reverse diffusion generation rollouts that the model executes. The dotted lines illustrate the bootstrapping feedback from the confidence model that is used to update the likelihood of the
Figure 1: Visual representation of the Confidence Bootstrapping training scheme. The dashed lines represent the reverse diffusion generation rollouts that the model executes. The dotted lines illustrate the bootstrapping feedback from the confidence model that is used to update the likelihood of the

실험 결과

연구 질문

  • RQ1기존 ML 기반 도킹 방법이 보지 못한 단백질 포켓 도메인으로 얼마나 잘 일반화하나요?
  • RQ2학습 데이터 및 모델 크기의 증가가 Blind docking 일반화에 어떤 영향을 미치나요?
  • RQ3합성 데이터 확장이 포켓 다양성을 넓히고 도킹 일반화를 개선할 수 있나요?
  • RQ4확신 기반 부트스트래핑이 보지 못한 단백질 클래스로의 확산 기반 도킹을 개선할 수 있나요?
  • RQ5더 크고 보강된 확산 모델이 DockGen 벤치마크에서 어떤 성능을 보이나요?

주요 결과

  • DockGen은 보지 못한 포켓에서 기존 ML 도킹 방식의 일반화 격차를 크게 보여줍니다.
  • DiffDock-L은 DockGen에서 Top-1 성공률을 7.1%에서 22.6%로 향상시키며 강력한 베이스라인을 능가하고 최첨단 결과를 달성합니다.
  • MOAD로 학습 데이터를 보강하고 van der Mer 영감을 받은 합성 리간드를 사용하면 포켓 다양성이 증가하고 성능이 소폭 향상됩니다.
  • 더 큰 모델 크기(예: 3000만 파라미터)가 데이터 보강과 결합될 때 일반화 이득을 제공합니다.
  • Confidence Bootstrapping은 DockGen-clusters에서 DiffDock-S의 성능을 9.8%에서 24.0%로 크게 높이고, 절반 이상 클러스터에서 30% 이상을 달성합니다.
  • DockGen 전반에서 확신이 강화된 확산 접근법은 높은 소모도에서 전통적 탐색 방법보다 우수한 성능을 보입니다.
Figure 2: A. An example of the superimposition of the pockets of two proteins in PDBBind, 1QXZ in pink and 5M4Q in cyan, that share a very similar binding pocket structure (a bound ligand is shown in red), but have only 22% sequence similarity. While sequence similarity splits would classify them in
Figure 2: A. An example of the superimposition of the pockets of two proteins in PDBBind, 1QXZ in pink and 5M4Q in cyan, that share a very similar binding pocket structure (a bound ligand is shown in red), but have only 22% sequence similarity. While sequence similarity splits would classify them in

더 나은 연구,지금 바로 시작하세요

논문 읽기부터 검토까지, 연구 시간을 획기적으로 줄여보세요.

카드 등록 없음 · 무료 플랜 제공

이 리뷰는 AI가 만들고, 인간 에디터가 검토했습니다.