Search bioRxiv⌕ Search

Biology subjects

Bastien, G. E.

Publications and source records attributed to Bastien, G. E..

2 recordsLinked to original sources

Virus-Host Interactions Predictor (VHIP): machine learning approach to resolve microbial virus-host interaction networks

Viruses of microbes are ubiquitous biological entities that reprogram their hosts metabolisms during infection in order to produce viral progeny, impacting the ecology and evolution of microbiomes with broad implications for human and environmental health. Advances in genome sequencing have led to the discovery of millions of novel viruses and an appreciation for the great diversity of viruses on Earth. Yet, with knowledge of only "who is there?" we fall short in our ability to infer the impacts of viruses on microbes at population, community, and ecosystem-scales. To do this, we need a more explicit understanding "who do they infect?" Here, we developed a novel machine learning model (ML), Virus-Host Interaction Predictor (VHIP), to predict virus-host interactions (infection/non-infection) from input virus and host genomes. This ML model was trained and tested on a high-value manually curated set of 8849 virus-host pairs and their corresponding sequence data. The resulting dataset, Virus Host Range network (VHRnet), is core to VHIP functionality. Each data point that underlies the VHIP training and testing represents a lab-tested virus-host pair in VHRnet, from which features of coevolution were computed. VHIP departs from existing virus-host prediction models in its ability to predict multiple interactions rather than predicting a single most likely host or host clade. As a result, VHIP is the first virus-host range prediction tool able to reconstruct the complexity of virus-host networks in natural systems. VHIP has an 87.8% accuracy rate at predicting interactions between virus-host pairs at the species level and can be applied to novel viral and host population genomes reconstructed from metagenomic datasets. Through the reconstruction of complete virus-host networks from novel data, VHIP allows for the integration of multilayer network theory into microbial ecology and opens new opportunities to study ecological complexity in microbial systems. Author summaryThe ecology and evolution of microbial communities are deeply influenced by viruses. Metagenomics analysis, the non-targeted sequencing of community genomes, has led to the discovery of millions of novel viruses. Yet, through the sequencing process, only DNA sequences are recovered, begging the question: which microbial hosts do those novel viruses infect? To address this question, we developed a computational tool to allow researchers to predict virus-host interactions from such sequence data. The power of this tool is its use of a high-value, manually curated set of 8849 lab-verified virus-host pairs and their corresponding sequence data. For each pair, we computed signals of coevolution to use as the predictive features in a machine learning model designed to predict interactions between viruses and hosts. The resulting model, Virus-Host Interaction Predictor (VHIP), has an accuracy of 87.8% and can be applied to novel viral and host genomes reconstructed from metagenomic datasets. Because the model considers all possible virus-host pairs, it can resolve complete virus-host interaction networks and supports a new avenue to apply network thinking to viral ecology.

bioinformatics↗

Benchmarking combined informatics approaches for virus discovery: Caution is needed when combining in silico identification methods

BackgroundThe identification of viruses from environmental metagenomic samples using informatics tools has offered critical insights in microbiome studies. However, it remains difficult for researchers to know for their specific study which tool(s) and settings are best suited to maximize capture of viruses while minimizing false positives. Studies are increasingly combining multiple tool outputs attempting to recover more viruses, but no combined approach has been benchmarked for accuracy. Here, we benchmarked 63 viral identification rulesets against mock metagenomes composed of publicly available viral, bacterial, archaeal, fungal, and protist sequences. These rulesets are based on combinations of four single-tool rules and two multi-tool tuning rules. We applied these rulesets to various aquatic metagenomes and filtering strategies to evaluate the impact of habitat and viral enrichment on individual and combined tool performance. We provide a packaged pipeline for researchers that want to replicate our process. ResultsWe found that combining rules increased viral recall, but at the expense of increased false positives. Six of the 63 combinations tested had equivalent accuracies to the highest one (MCC=0.77, padj [≥] 0.05). All of the six high accuracy rulesets included VirSorter2, five included our "tuning removal" rule, and no high performing rulesets used more than four of our six rules. DeepVirFinder, VIBRANT, and VirSorter were each found once in these high accuracy rulesets, but never in combination with each other. Our validation suggests that the MCC plateau at 0.77 is caused by inaccurate labeling of the data that viral identification tools rely on for training and validation. In the aquatic metagenomes, our "highest MCC" ruleset identified a higher proportion of viral sequences in the virus-enriched samples (44-46%) than the non-enriched, cellular metagenomes (7-19%). ConclusionWhile improved algorithms may lead to more accurate viral identification tools, this should be done in tandem with curating accurately labeled viral gene and sequence databases. For most applications, we recommend the use of the ruleset that uses VirSorter2 and our empirically derived tuning removal rule. By providing a rigorous overview of the behavior of in silico viral identification strategies, our findings guide the use of existing viral identification tools and offer a blueprint for feature engineering of new tools that will lead to higher-confidence viral discovery in microbiome studies.

bioinformatics↗