Search bioRxiv⌕ Search

Biology subjects

Schiffer, A. M.

Publications and source records attributed to Schiffer, A. M..

2 recordsLinked to original sources

A comparison of short- and long-read whole genome sequencing for microbial pathogen epidemiology

Whole genome sequencing provides the highest resolution for characterizing pathogen evolution, epidemiology, and diagnostics. Genome assemblies contain information on the identity and potential phenotypes of a pathogen. Likewise, variant calling can inform on transmission patterns and evolutionary relationships. Recent improvements in Oxford Nanopore long-read sequencing have made its use attractive for genomic epidemiology. However, the accuracy and optimal strategy for analysis of Nanopore reads remains to be determined. We compared the use of Illumina short reads and Oxford Nanopore long reads for genome assembly and variant calling of phytopathogenic bacteria. We generated short- and long-read datasets for diverse phytopathogenic Agrobacterium strains. We then analyzed these data using multiple pipelines designed for either short or long reads and compared the results. We found that assemblies made from long reads were more complete than those made from short-read data and contained few sequence errors. Variant calling pipelines differed in their ability to accurately call variants and infer genotypes from long reads. Results suggest that computationally fragmenting long reads can improve the accuracy of variant calling in population-level studies. Using fragmented long reads, pipelines designed for short reads were more accurate at recovering genotypes than pipelines designed for long reads. Further, short- and long-read datasets can be analyzed together with the same pipelines. These findings show that Oxford Nanopore sequencing is accurate and can be sufficient for microbial pathogen genomics and epidemiology. Ultimately, this enhances the ability of researchers and clinicians to understand and mitigate the spread of pathogens. ImportanceGenome assembly and variant calling are important steps in microbial population studies and epidemiology. Most variant calling and genotyping pipelines are designed for Illumina short sequencing reads. Oxford Nanopore Technology long-read sequencing results in more complete genome assemblies but has historically been of lower quality. Here, we show that Nanopore long reads are now of sufficient quality for bacterial whole genome assembly and epidemiology. We benchmarked the accuracy of multiple variant-calling pipelines with short and long reads. Using an optimized variant calling approach, variant calls and genotypes inferred from long reads are as accurate as those inferred from short reads. Importantly, we found that gold-standard variant calling pipelines designed for short reads are also accurate with long reads when long reads are first fragmented into shorter sequences. This finding allows researchers to incorporate the advantages of Nanopore sequencing for genome assembly, while maintaining high accuracy for epidemiology and population analysis.

genomics↗

Beav: A bacterial genome and mobile element annotation pipeline

Comprehensive and accurate genome annotation is crucial for inferring the predicted functions of an organism. Numerous tools exist to annotate genes, gene clusters, mobile genetic elements, and other diverse features. However, these tools and pipelines can be difficult to install and run, be specialized for a particular element or feature, or lack annotations for larger elements that provide important genomic context. Integrating results across analyses is also important for understanding gene function. To address these challenges, we present the Beav annotation pipeline. Beav is a command-line tool that automates the annotation of bacterial genome sequences, mobile genetic elements, molecular systems and gene clusters, key regulatory features, and other elements. Beav uses existing tools in addition to custom models, scripts, and databases to annotate diverse elements, systems, and sequence features. Custom databases for plant-associated microbes are incorporated to improve annotation of key virulence and symbiosis genes in agriculturally important pathogens and mutualists. Beav includes an optional Agrobacterium-specific pipeline that identifies and classifies oncogenic plasmids and annotates plasmid-specific features. Following the completion of all analyses, annotations are consolidated to produce a single comprehensive output. Finally, Beav generates publication-quality genome and plasmid maps. Beav is on Bioconda and is available for download at https://github.com/weisberglab/beav.

bioinformatics↗