[Paper Review] Pandemic Drugs at Pandemic Speed: Infrastructure for Accelerating COVID-19 Drug Discovery with Hybrid Machine Learning- and Physics-based Simulations on High Performance Computers
This paper presents a hybrid machine learning (ML) and physics-based (PB) simulation workflow accelerated by high-performance computing (HPC) to identify antiviral drug candidates for SARS-CoV-2 at unprecedented speed. By integrating thermodynamic integration with enhanced sampling (TIES) for accurate binding affinity prediction and ML models for efficient chemical space exploration, the framework enables high-throughput screening of millions of compounds, successfully identifying lead compounds for four viral targets with quantitative free energy predictions.
The race to meet the challenges of the global pandemic has served as a reminder that the existing drug discovery process is expensive, inefficient and slow. There is a major bottleneck screening the vast number of potential small molecules to shortlist lead compounds for antiviral drug development. New opportunities to accelerate drug discovery lie at the interface between machine learning methods, in this case developed for linear accelerators, and physics-based methods. The two in silico methods, each have their own advantages and limitations which, interestingly, complement each other. Here, we present an innovative infrastructural development that combines both approaches to accelerate drug discovery. The scale of the potential resulting workflow is such that it is dependent on supercomputing to achieve extremely high throughput. We have demonstrated the viability of this workflow for the study of inhibitors for four COVID-19 target proteins and our ability to perform the required large-scale calculations to identify lead antiviral compounds through repurposing on a variety of supercomputers.
Motivation & Objective
- Address the bottleneck in traditional drug discovery, which takes ~10 years and $1–3 billion, by accelerating the identification of antiviral compounds for SARS-CoV-2.
- Overcome the limitations of standalone machine learning and physics-based methods by integrating their complementary strengths: ML for speed and PB for accuracy.
- Develop a scalable, high-throughput computational infrastructure capable of handling large-scale simulations across diverse supercomputers.
- Enable rapid repurposing of existing drugs by predicting binding affinities of ligands with key SARS-CoV-2 target proteins with high precision.
- Create an iterative feedback loop between ML and PB simulations to refine predictions and guide chemical modifications toward optimal lead compounds.
Proposed method
- Employ thermodynamic integration with enhanced sampling (TIES) for accurate, reproducible, and scalable free energy calculations of protein-ligand binding affinities.
- Use ensemble simulations to reduce variability from chaotic initial conditions and improve statistical reliability of binding affinity predictions.
- Integrate TIES-predicted binding affinities as training data to iteratively improve machine learning models for predicting ligand efficacy.
- Implement a multi-stage filtering workflow: initial ML screening followed by high-accuracy TIES validation on top candidates.
- Leverage high-performance computing (HPC) infrastructure across multiple supercomputers, including SuperMUC-NG, OLCF, and TACC, to scale simulations.
- Use a dedicated workflow manager (e.g., RCT) to orchestrate heterogeneous, compute-intensive tasks across distributed HPC resources.
Experimental results
Research questions
- RQ1Can a hybrid ML-PB simulation workflow significantly accelerate the identification of lead compounds for SARS-CoV-2 drug repurposing?
- RQ2How accurately can TIES-based free energy calculations predict binding affinities for diverse ligands targeting SARS-CoV-2 proteins?
- RQ3To what extent does integrating PB simulation data into ML models improve the predictive accuracy and convergence speed in chemical space exploration?
- RQ4Can this hybrid workflow be efficiently scaled across multiple heterogeneous HPC systems to achieve pandemic-speed drug discovery?
- RQ5What structural transformations in ligands lead to favorable or unfavorable changes in binding affinity, and can these be reliably predicted?
Key findings
- The hybrid ML-PB workflow successfully identified lead antiviral compounds for four SARS-CoV-2 target proteins using large-scale simulations on multiple supercomputers.
- TIES simulations achieved high accuracy and precision in binding affinity predictions, with uncertainties as low as ±0.44 kcal/mol for certain transformations.
- For the ADRP target, 12 out of 19 ligand transformations resulted in ΔΔ𝐺 > +1 kcal/mol, indicating unfavorable binding changes, while seven had ΔΔ𝐺 ≈ 0, suggesting no significant effect.
- The transformation a42-a43 showed the largest unfavorable change with ΔΔ𝐺 = 4.62 ± 0.82 kcal/mol, indicating a strong destabilization of binding.
- The workflow enabled rapid, iterative refinement of ML models using PB-derived energetic data, accelerating convergence toward promising chemical space regions.
- The entire pipeline was successfully deployed across diverse HPC systems, including SuperMUC-NG, OLCF, and TACC, demonstrating portability and scalability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.