[Paper Review] Network embedding unveils the hidden interactions in the mammalian virome
This study introduces a novel LF-SVD method combining linear filtering and singular value decomposition to predict hidden and unrealized host-virus interactions in the mammalian virome. By leveraging network structure and low-rank imputation, the approach recovers biologically plausible associations, reveals the Amazon as a global hotspot of unsampled virome diversity, and enhances zoonotic risk prediction using graph-embedded viral features.
At most 1-2% of the global virome has been sampled to date. Recent work has shown that predicting which host-virus interactions are possible but undiscovered or unrealized is, fundamentally, a network science problem. Here, we develop a novel method that combines a coarse recommender system (Linear Filtering; LF) with an imputation algorithm based on low-rank graph embedding (Singular Value Decomposition; SVD) to infer host-virus associations. This combination of techniques results in informed initial guesses based on directly measurable network properties (density, degree distribution) that are refined through SVD (which is able to leverage emerging features). Using this method, we recovered highly plausible undiscovered interactions with a strong signal of viral coevolutionary history, and revealed a global hotspot of unusually unique but unsampled (or unrealized) host-virus interactions in the Amazon rainforest. We develop several tests for quantifying the bias and realism of these predictions, and show that the LF-SVD method is robust in each aspect. We finally show that graph embedding of the imputed network can be used to improve predictions of human infection from viral genome features, showing that the global structure of the mammal-virus network provides additional insights into human disease emergence.
Motivation & Objective
- To address the critical gap in understanding the global virome, where less than 1-2% has been sampled to date.
- To predict biologically plausible but unrealized host-virus interactions in a high-sparsity, partially observed network.
- To reduce sampling bias in network predictions while preserving coevolutionary signals in host-virus associations.
- To improve zoonotic risk prediction by integrating global network structure into viral genome-based models.
- To test whether graph-embedded features from imputed networks enhance the accuracy of human infection risk prediction for mammalian viruses.
Proposed method
- The LF-SVD method uses linear filtering (LF) to generate initial interaction probabilities based on network properties: in-degree, out-degree, and connectance.
- Singular value decomposition (SVD) is applied to a low-rank approximation of the network adjacency matrix to impute missing or unrealized interactions.
- The model is tuned to maximize information use by selecting an optimal matrix rank (rank 12 in the best-performing model) and weighting connectance over degree to reduce sampling bias.
- Graph embedding is performed on both observed and imputed networks using a random dot product graph decomposition to extract latent traits (12 dimensions) from the left singular subspace.
- These latent features are combined with viral genome composition metrics to train gradient boosted tree models for zoonotic potential prediction.
- To prevent data leakage, human hosts and viruses linked exclusively to humans are excluded during embedding generation, and model performance is validated using repeated 1000-fold train-calibrate-test splits.
Experimental results
Research questions
- RQ1Can a network-based method infer biologically plausible but unrealized host-virus interactions in a sparsely sampled virome?
- RQ2Does the LF-SVD method reduce the influence of sampling bias while preserving evolutionary signals in host-virus associations?
- RQ3Is the Amazon rainforest a global hotspot for undiscovered host-virus interactions, as suggested by network imputation?
- RQ4Can graph-embedded features derived from an imputed network improve the prediction of zoonotic potential from viral genome composition alone?
- RQ5Does incorporating global network structure enhance the accuracy of models predicting which mammalian viruses can infect humans?
Key findings
- The LF-SVD method achieved a ROC-AUC of 0.84 in predicting known host-virus interactions, demonstrating strong predictive performance on a partially observed network.
- Imputation increased the number of predicted interactions by approximately 15-fold, indicating a vast reservoir of previously undetected or unrealized host-virus associations.
- The Amazon rainforest emerged as a global hotspot of unique, unsampled host-virus interactions, with LCBD analysis showing higher uniqueness in viral community composition than other regions.
- The model significantly reduced the influence of passive sampling bias: citation counts had a weaker effect on predicting viral richness after imputation, indicating reduced bias in predictions.
- Graph embeddings derived from the imputed network improved zoonotic risk prediction, with the best model achieving higher ROC-AUC when combining viral genome features and imputed network embeddings.
- The final model, trained on genome composition and imputed network embeddings, predicted human infection probabilities for 612 viruses with improved accuracy through ensemble averaging across top-performing models.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.