[Paper Review] Targeting SARS-CoV-2 with AI- and HPC-enabled Lead Generation: A First Data Release
This paper presents a large-scale, AI- and HPC-powered data release of over 4.2 billion molecules enriched with molecular fingerprints, 2D images, and 2D/3D descriptors to accelerate SARS-CoV-2 drug discovery. The dataset, generated using high-performance computing and made publicly accessible via Globus, enables rapid machine learning screening for repurposed or novel antiviral compounds.
Researchers across the globe are seeking to rapidly repurpose existing drugs or discover new drugs to counter the the novel coronavirus disease (COVID-19) caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). One promising approach is to train machine learning (ML) and artificial intelligence (AI) tools to screen large numbers of small molecules. As a contribution to that effort, we are aggregating numerous small molecules from a variety of sources, using high-performance computing (HPC) to computer diverse properties of those molecules, using the computed properties to train ML/AI models, and then using the resulting models for screening. In this first data release, we make available 23 datasets collected from community sources representing over 4.2 B molecules enriched with pre-computed: 1) molecular fingerprints to aid similarity searches, 2) 2D images of molecules to enable exploration and application of image-based deep learning methods, and 3) 2D and 3D molecular descriptors to speed development of machine learning models. This data release encompasses structural information on the 4.2 B molecules and 60 TB of pre-computed data. Future releases will expand the data to include more detailed molecular simulations, computed models, and other products.
Motivation & Objective
- To accelerate drug discovery for SARS-CoV-2 by enabling large-scale, AI-driven screening of small molecules.
- To overcome computational barriers by pre-computing molecular descriptors and fingerprints using high-performance computing (HPC).
- To provide a unified, accessible data resource that integrates diverse molecular datasets from public and proprietary sources.
- To support the development of machine learning models for identifying promising drug candidates for further experimental validation.
- To lay the foundation for future data releases including molecular conformers, docking simulations, and NLP-extracted drug candidates.
Proposed method
- Aggregated 23 molecular datasets from public and internal sources, including drug banks, antiviral libraries, and benchmark decoy sets.
- Used high-performance computing (HPC) to pre-compute 2D and 3D molecular descriptors, fingerprints, and 2D molecular images for all 4.2 billion molecules.
- Represented molecules using canonical SMILES strings and enriched them with structural and physicochemical properties to support ML model training.
- Hosted the data on Globus endpoints to enable secure, high-speed access and transfer via web interface, REST API, or Python SDK.
- Provided a unified data access layer using Globus, enabling researchers to browse, download, and transfer large-scale datasets efficiently.
- Designed the data pipeline to support future extensions, including natural language processing of scientific literature and molecular docking simulations.
Experimental results
Research questions
- RQ1Can large-scale, HPC-accelerated computation of molecular descriptors and fingerprints significantly reduce the time and resource burden for AI-based drug screening?
- RQ2How can diverse, heterogeneous molecular datasets be unified into a single, accessible, and computationally enriched resource for SARS-CoV-2 research?
- RQ3To what extent can pre-computed molecular features improve the accuracy and speed of machine learning models in identifying potential antiviral compounds?
- RQ4How can high-performance data transfer infrastructure like Globus enhance collaboration and reproducibility in large-scale computational drug discovery?
- RQ5What role can AI and HPC play in accelerating the identification of repurposed or novel small molecules for SARS-CoV-2?
Key findings
- The first data release includes 23 datasets representing over 4.2 billion unique molecules with pre-computed molecular fingerprints, 2D images, and 2D/3D descriptors.
- A total of 60 terabytes of enriched molecular data were generated and made publicly available through high-performance data transfer infrastructure.
- The dataset integrates diverse sources including DrugBank, Enamine, CAS COVID-19 antiviral candidates, DUDE decoys, and QM9, enabling broad applicability for model training.
- The data are accessible via a web interface, Globus endpoints, and a Python SDK, supporting scalable data access and transfer to local or HPC systems.
- The release enables researchers to bypass computationally expensive descriptor calculations and directly train or deploy ML models for virtual screening.
- Future releases will expand the dataset to include molecular conformers, results from natural language processing of scientific literature, and pre-screened candidates from molecular docking simulations.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.