Search bioRxiv⌕ Search

Biology subjects

Mayne, R. M.

Publications and source records attributed to Mayne, R. M..

2 recordsLinked to original sources

GRAViTy-V2: a grounded viral taxonomy application

Taxonomic classification of viruses is essential for understanding their evolution and therefore their distribution, host interactions and pathogenic mechanisms. Classification methodologies usually rely on comparison of aligned sequence motifs in conserved genes, by genome organisation and gene complements, and at lower taxonomic ranks such as genus and species, through genome sequence identities. Building on our previous classification framework based on a novel whole-genome analysis method, we here describe Genome Relationships Applied to Viral Taxonomy Version 2 (GRAViTy-V2), which encompasses a greatly expanded range of features and numerous optimisations, packaged as an application that may be used as an alignment-free general-purpose virus classification tool. Using 28 datasets derived from the International Society on Taxonomy of Viruses 2022 taxonomy proposals, GRAViTy-V2 output was compared against human expert-curated classifications used for assignments in the 2023 round of ICTV taxonomy changes. GRAViTy-V2 produced taxonomies equivalent to manually-curated versions down to the family level and in almost all cases, to genus and species levels. However, discrepancies with our results primarily arose through various human and automated sequence annotation errors and erroneous annotations of coding sequences used in their original classification. Analysis times ranged from 1-506 min (median 3.59) on datasets with 17-1004 genomes and mean genome length of 3,000-1,000,000 bases, on a standard consumer-grade laptop. We discuss how the output from GRAViTY-V2 outputs allows for a full analysis of why taxonomic classifications were proposed, the value of the program for quality control of genetic comparisons, and how to optimise the speed of classification through proper use of GRAViTy-V2s workflow management system.

bioinformatics↗

Castanet: a pipeline for rapid analysis of targeted multi-pathogen genomic data

MotivationTarget enrichment strategies generate genomic data from multiple pathogens in a single process, greatly improving sensitivity over metagenomic sequencing and enabling cost-effective, high throughput surveillance and clinical applications. However, uptake by research and clinical laboratories is constrained by an absence of computational tools that are specifically designed for the analysis of multi-pathogen enrichment sequence data. Here we present the Castanet pipeline: an analysis pipeline for end-to-end processing and consensus sequence generation for use with multi-pathogen enrichment sequencing data. Castanet is designed to work with short-read data produced by existing targeted enrichment strategies, but can be readily deployed on any BAM file generated by another methodology. It is packaged with usability features, including graphical interface and installer script. ResultsIn addition to genome reconstruction, Castanet reports method-specific metrics that enable quantification of capture efficiency, estimation of pathogen load, differentiation of low-level positives from contamination, and assessment of sequencing quality. Castanet can be used as a traditional end-to-end pipeline for consensus generation, but its strength lies in the ability to process a flexible, pre-defined set of pathogens of interest directly from multi-pathogen enrichment experiments. In our tests, Castanet consensus sequences were accurate reconstructions of reference sequences, including in instances where multiple strains of the same pathogen were present. Castanet performs effectively on standard laptop computers and can process the entire output of a 96-sample enrichment sequencing run (50M reads) using a single batch process command, in < 2 h. Availability and ImplementationSource code freely available under GPL-3 license at https://github.com/MultipathogenGenomics/castanet, implemented in Python 3.10 and supported in Ubuntu Linux 22.04 and other Bash-like environments. The data for this study have been deposited in the European Nucleotide Archive (ENA) at EMBL-EBI under accession number PRJEB77004.

bioinformatics↗