Search bioRxiv⌕ Search

Biology subjects

Urhan, A.

Publications and source records attributed to Urhan, A..

3 recordsLinked to original sources

Global diversity of enterococci and description of 18 novel species

Enterococci are commensal gut microbes of most land animals. They diversified over hundreds of millions of years adapting to evolving hosts and host diets. Of over 60 known enterococcal species, Enterococcus faecalis and E. faecium uniquely emerged in the antibiotic era among leading causes of multidrug resistant hospital-associated infection. The basis for the association of particular enterococcal species with a host is largely unknown. To begin deciphering enterococcal species traits that drive host association, and to assess the pool of Enterococcus-adapted genes from which known facile gene exchangers such as E. faecalis and E. faecium may draw, we collected 886 enterococcal strains from nearly 1,000 specimens representing widely diverse hosts, ecologies and geographies. This provided data on the global occurrence and host associations of known species, identifying 18 new species in the process expanding genus diversity by >25%. The novel species harbor diverse genes associated with toxins, detoxification, and resource acquisition. E. faecalis and E. faecium were isolated from a wide diversity of hosts highlighting their generalist properties, whereas most other species exhibited more restricted distributions indicative of specialized host associations. The expanded species diversity permitted the Enterococcus genus phylogeny to be viewed with unprecedented resolution, allowing features to be identified that distinguish its four deeply rooted clades as well as genes associated with range expansion, such as B-vitamin biosynthesis and flagellar motility. Collectively, this work provides an unprecedentedly broad and deep view of the genus Enterococcus, potential threats to human health, and new insights into its evolution. SIGNIFICANCEEnterococci, host-associated microbes that are now leading drug-resistant hospital pathogens, arose as animals colonized land over 400 million years ago. To globally assess the diversity of enterococci now associated with land animals, we collected 886 enterococcal specimens from a wide range of geographies and ecologies, ranging from urban environments to remote areas generally inaccessible to humans. Species determination and genome analysis revealed host associations from generalists to specialists, and identified 18 new species, increasing the genus by over 25%. This added diversity provided greater resolution of the genus clade structure, identifying new features associated with species radiations. Moreover, the high rate of new species discovery shows that tremendous genetic diversity in Enterococcus remains to be discovered.

microbiology↗

SAP: Synteny-aware gene function prediction for bacteria using protein embeddings

MotivationToday, we know the function of only a small fraction of the protein sequences predicted from genomic data. This problem is even more salient for bacteria, which represent some of the most phylogenetically and metabolically diverse taxa on Earth. This low rate of bacterial gene annotation is compounded by the fact that most function prediction algorithms have focused on eukaryotes, and conventional annotation approaches rely on the presence of similar sequences in existing databases. However, often there are no such sequences for novel bacterial proteins. Thus, we need improved gene function prediction methods tailored for prokaryotes. Recently, transformer-based language models - adopted from the natural language processing field - have been used to obtain new representations of proteins, to replace amino acid sequences. These representations, referred to as protein embeddings, have shown promise for improving annotation of eukaryotes, but there have been only limited applications on bacterial genomes. ResultsTo predict gene functions in bacteria, we developed SAP, a novel synteny-aware gene function prediction tool based on protein embeddings from state-of-the-art protein language models. SAP also leverages the unique operon structure of bacteria through conserved synteny. SAP outperformed both conventional sequence-based annotation methods and state-of-the-art methods on multiple bacterial species, including for distant homolog detection, where the sequence similarity to the proteins in the training set was as low as 40%. Using SAP to identify gene functions across diverse enterococci, of which some species are major clinical threats, we identified 11 previously unrecognized putative novel toxins, with potential significance to human and animal health. Availabilityhttps://github.com/AbeelLab/sap Contactt.abeel@tudelft.nl Supplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics↗

HAT: Haplotype Assembly Tool using short and error-prone long reads

MotivationHaplotypes are the set of alleles cooccurring on a single chromosome and inherited together to the next generation. Because a monoploid reference genome loses this co-occurrence information, its usability is limited to associate phenotypes with allelic combinations of genotypes. Therefore, methods to reconstruct the complete haplotypes from DNA sequencing data are crucial. Recently, several attempts have been made at haplotype reconstructions, but significant limitations remain. High-quality continuous haplotypes cannot be created reliably, particularly when there are few differences between the homologous chromosomes. ResultsHere, we introduce HAT, a haplotype assembly tool that exploits short and long reads along with a reference genome to reconstruct haplotypes. HAT tries to take advantage of the accuracy of short reads and the length of the long reads to reconstruct haplotypes. We tested HAT on the aneuploid yeast strain Saccharomyces pastorianus CBS1483 and multiple simulated polyploid data sets of the same strain, showing that it outperforms existing tools. Availabilityhttps://github.com/AbeelLab/hat/ Contactt.abeel@tudelft.nl

bioinformatics↗