Search bioRxivSearch

Biology subjects

Mussig, A. J.

Publications and source records attributed to Mussig, A. J..

2 recordsLinked to original sources

A rank-normalized archaeal taxonomy based on genome phylogeny resolves widespread incomplete and uneven classifications

An increasing wealth of genomic data from cultured and uncultured microorganisms provides the opportunity to develop a systematic taxonomy based on evolutionary relationships. Here we propose a standardized archaeal taxonomy, as part of the Genome Taxonomy Database (GTDB), derived from a 122 concatenated protein phylogeny that resolves polyphyletic groups and normalizes ranks based on relative evolutionary divergence (RED). The resulting archaeal taxonomy is stable under a range of phylogenetic variables, including marker genes, inference methods, corrections for rate heterogeneity and compositional bias, tree rooting scenarios, and expansion of the genome database. Rank normalization was shown to robustly correct for substitution rates varying up to 30-fold using simulated datasets. Taxonomic curation follows the rules of the International Code of Nomenclature of Prokaryotes (ICNP) while taking into account proposals to formally recognise the rank of phylum and to use genome sequences as type material. The taxonomy is based on 2,392 quality screened archaeal genomes, the great majority of which (93.3%) required one or more changes to their existing taxonomy, mostly as a result of incomplete classification. In total, 16 archaeal phyla are described, including reclassification of three major monophyletic units from the former Euryarchaeota and one phylum resulting from uniting the TACK superphylum into a single phylum. The taxonomy is publicly available at the GTDB website (https://gtdb.ecogenomic.org).

microbiology

Selection of representative genomes for 24,706 bacterial and archaeal species clusters provide a complete genome-based taxonomy

We recently introduced the Genome Taxonomy Database (GTDB), a phylogenetically consistent, genome-based taxonomy providing rank normalized classifications for nearly 150,000 genomes from domain to genus. However, nearly 40% of the genomes used to infer the GTDB reference tree lack a species name, reflecting the large number of genomes in public repositories without complete taxonomic assignments. Here we address this limitation by proposing 24,706 species clusters which encompass all publicly available bacterial and archaeal genomes when using commonly accepted average nucleotide identity (ANI) criteria for circumscribing species. In contrast to previous ANI studies, we selected a single representative genome to serve as the nomenclatural type for circumscribing each species with type strains used where available. We complemented the 8,792 species clusters with validly or effectively published names with 15,914 de novo species clusters in order to assign placeholder names to the growing number of genomes from uncultivated species. This provides the first complete domain to species taxonomic framework which will improve communication of scientific results.

microbiology