[Paper Review] Pre-Training on Large-Scale Generated Docking Conformations with HelixDock to Unlock the Potential of Protein-ligand Structure Prediction Models
This paper proposes HelixDock, a geometry-aware SE(3)-equivariant deep learning model pre-trained on 100 million protein-ligand docking conformations generated by physics-based tools, then fine-tuned on limited experimental data. The method achieves a 40%+ improvement in RMSD over state-of-the-art models, demonstrating superior accuracy and robustness in protein-ligand structure prediction by leveraging large-scale generated data to embed physical knowledge.
Protein-ligand structure prediction is an essential task in drug discovery, predicting the binding interactions between small molecules (ligands) and target proteins (receptors). Recent advances have incorporated deep learning techniques to improve the accuracy of protein-ligand structure prediction. Nevertheless, the experimental validation of docking conformations remains costly, it raises concerns regarding the generalizability of these deep learning-based methods due to the limited training data. In this work, we show that by pre-training on a large-scale docking conformation generated by traditional physics-based docking tools and then fine-tuning with a limited set of experimentally validated receptor-ligand complexes, we can obtain a protein-ligand structure prediction model with outstanding performance. Specifically, this process involved the generation of 100 million docking conformations for protein-ligand pairings, an endeavor consuming roughly 1 million CPU core days. The proposed model, HelixDock, aims to acquire the physical knowledge encapsulated by the physics-based docking tools during the pre-training phase. HelixDock has been rigorously benchmarked against both physics-based and deep learning-based baselines, demonstrating its exceptional precision and robust transferability in predicting binding confirmation. In addition, our investigation reveals the scaling laws governing pre-trained protein-ligand structure prediction models, indicating a consistent enhancement in performance with increases in model parameters and the volume of pre-training data. Moreover, we applied HelixDock to several drug discovery-related tasks to validate its practical utility. HelixDock demonstrates outstanding capabilities on both cross-docking and structure-based virtual screening benchmarks.
Motivation & Objective
- To address the data scarcity and limited generalization of deep learning-based protein-ligand structure prediction models.
- To improve prediction accuracy by pre-training on large-scale, physically plausible docking conformations generated by traditional tools.
- To investigate the scaling laws governing pre-trained protein-ligand structure prediction models.
- To enhance model robustness on challenging and novel protein-ligand complexes, including those relevant to SARS-CoV-2.
- To demonstrate the feasibility and superiority of combining physics-based conformation generation with deep learning for AI-driven drug discovery.
Proposed method
- Pre-training a SE(3)-equivariant neural network on 100 million protein-ligand docking conformations generated by physics-based docking tools using HelixDock.
- Utilizing a large-scale, diverse dataset of generated conformations to embed physical knowledge from force fields and scoring functions.
- Fine-tuning the pre-trained model on a limited set of experimentally validated receptor-ligand complexes to adapt to real binding geometries.
- Employing structure alignment strategies and interaction fingerprint similarity to evaluate model generalization and test set diversity.
- Applying fpocket to compute pocket volumes and Tanimoto similarity on PLEC fingerprints to assess structural and chemical similarity to training data.
- Benchmarking against physics-based and deep learning baselines on PDBBind and PDB-COVID19 datasets to evaluate RMSD, affinity prediction, and robustness.
Experimental results
Research questions
- RQ1Can pre-training on large-scale, generated docking conformations significantly improve the accuracy of deep learning-based protein-ligand structure prediction models?
- RQ2How does the performance of the model scale with increases in model size and pre-training data volume?
- RQ3Does pre-training on physics-based generated data enhance generalization to novel protein-ligand complexes, especially those with limited experimental data?
- RQ4How does the model perform on challenging targets, such as SARS-CoV-2-related proteins, compared to state-of-the-art methods?
- RQ5What is the impact of physical knowledge embedded in generated conformations on downstream prediction accuracy and robustness?
Key findings
- HelixDock achieves a 40%+ improvement in RMSD over its closest competitor on standard benchmark datasets, with an RMSD of 2.06 Å on the PDBBind core set.
- The model demonstrates superior performance on the PDB-COVID19 dataset, outperforming baselines on 10 out of 11 primary SARS-CoV-2 protein targets.
- HelixDock achieves an affinity prediction RMSE of 1.159 and Pearson R of 0.844 on the PDBBind core set, outperforming all baselines.
- The model exhibits consistent performance gains with increasing model size and pre-training data volume, confirming observed scaling laws in protein-ligand structure prediction.
- Pre-training on generated data significantly enhances generalization, as evidenced by improved performance on diverse and challenging targets, including those dissimilar to training data.
- The ablation study shows that HelixDock without pre-training performs poorly (RMSE = 1.470), confirming the critical role of pre-training on large-scale generated conformations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.