Search bioRxiv⌕ Search

Biology subjects

Mitchell, J. A. M.

Publications and source records attributed to Mitchell, J. A. M..

2 recordsLinked to original sources

Models trained with noisy genomes extend bacterial phenotype prediction into deep time

Predicting phenotype from genotype in extant organisms is increasingly tractable through the accumulation of genome sequences and the development of machine-learning algorithms. Here we show that machine learning can be applied to reconstructed ancestral gene content, extending these predictions into the past. We trained models on a diverse set of bacterial phenotypes and found that introducing noise into gene content profiles allows predictions to generalize over larger evolutionary distances. For phenotypes with signal spread across many genes - such as metabolic oxygen use, cell envelope architecture and optimal growth temperature - noise augmentation extends resolution back to the root of the bacterial domain, while for other phenotypes - including GC content and sporulation - the range remains more limited. We therefore conclude that the last bacterial common ancestor (LBCA) was likely an anaerobic, double-membraned, and moderately thermophilic bacterium (46-75{degrees}C). Moreover, this work provides a general approach for learning about the genomic basis of phenotypes and drawing inferences about their early evolution.

evolutionary biology↗

SingleM and Sandpiper: Robust microbial taxonomic profiles from metagenomic data

Determining the taxonomy and relative abundance of microorganisms in metagenomic data is a foundational problem in microbial ecology. To address the limitations of existing approaches, we developed SingleM, which estimates community composition using conserved regions within universal marker genes. SingleM accurately profiles complex communities of known microbial species, and is the only tool that detects species without genomic representation, even those representing novel phyla. Given SingleMs computational efficiency, we applied it to 248,559 publicly available metagenomes and show that the vast majority of samples from marine, freshwater, sediment and soil environments are dominated by novel species lacking genomic representation (median relative abundance 75.0%). SingleM also provides a way to identify metagenomes for the recovery of novel metagenome-assembled genomes from lineages of interest, and can incorporate user-recovered genomes into its reference database to improve profiling resolution. Quantifying the full diversity of Bacteria and Archaea in metagenomic data shows that microbial genome databases are far from saturated.

microbiology↗