Search bioRxivSearch

Biology subjects

Sahl, J. W.

Publications and source records attributed to Sahl, J. W..

2 recordsLinked to original sources

Botulinum-neurotoxin-like sequences identified from an Enterococcus sp. genome assembly

Botulinum neurotoxins (BoNTs) are produced by diverse members of the Clostridia and result in a flaccid paralysis known as botulism. Exploring the diversity of BoNTs is important for the development of therapeutics and antitoxins. Here we describe a novel, bont-like gene cluster identified in a draft genome assembly for Enterococcus sp. 3G1_DIV0629 by querying publicly available genomic databases. The bont-like gene is found in a gene cluster similar to known bont gene clusters. Protease and binding motifs conserved in known BoNT proteins are present in the newly identified BoNT-like protein; however, it is currently unknown if the BoNT-like protein described here is capable of targeting neuronal cells resulting in botulism.

genomics

Bacterial genome reduction as a result of short read sequence assembly

High-throughput comparative genomics has changed our view of bacterial evolution and relatedness. Many genomic comparisons, especially those regarding the accessory genome that is variably conserved across strains in a species, are performed using assembled genomes. For completed genomes, an assumption is made that the entire genome was incorporated into the genome assembly, while for draft assemblies, often constructed from short sequence reads, an assumption is made that genome assembly is an approximation of the entire genome. To understand the potential effects of short read assemblies on the estimation of the complete genome, we downloaded all completed bacterial genomes from GenBank, simulated short reads, assembled the simulated short reads and compared the resulting assembly to the completed assembly. Although most simulated assemblies demonstrated little reduction, others were reduced by as much as 25%, which was correlated with the repeat structure of the genome. A comparative analysis of lost coding region sequences demonstrated that up to 48 CDSs or up to ~112,000 bases of coding region sequence, were missing from some draft assemblies compared to their finished counterparts. Although this effect was observed to some extent in 32% of genomes, only minimal effects were observed on pan-genome statistics when using simulated draft genome assemblies. The benefits and limitations of using draft genome assemblies should be fully realized before interpreting data from assembly-based comparative analyses.

genomics