Search bioRxiv⌕ Search

Biology subjects

Münch, P. C.

Publications and source records attributed to Münch, P. C..

3 recordsLinked to original sources

Enhanced multi-omic viral profiling from microbial community sequencing with BAQLaVa

Viruses are crucial components of microbial communities, both phage that infect bacterial community members as well as pathogenic and other eukaryotic viruses. However, they remain unobserved by most current technologies, due to combinations of experimental and analytical factors. To address the latter, we developed the BAQLaVa algorithm for high-resolution profiling of >120,000 viral species (viral genome bins, VGBs) via reference-based metagenome (MGX) or metatranscriptome (MTX) alignment to complementary nucleotide markers and proteome sets. In comprehensive benchmarking, BAQLaVa substantially outperformed alternatives, achieving species-level recall and precision regularly over 90%. We applied BAQLaVa to MGX and MTX samples from the HMP2 IBDMDB cohort to identify previously undescribed viral perturbations in inflammatory bowel diseases. Most notably, virome diversity was reduced in tandem with bacterial diversity during inflammation, in contrast to previous findings based on a narrower range of viral detection. A subset of viruses were enriched during IBD, associated with carriage of abortive infection anti-defense systems such as AbiL and PD-{lambda}-2, as well as genes involved in the regulation of lysogeny. Leveraging the corresponding viral profiles, we also inferred phage-host relationships using scalable co-occurrence and covariation signals, even in the absence of host references or genome annotations. By enabling high sensitivity and specificity viral profiling from metagenomes or metatranscriptomes, BAQLaVa provides a scalable framework for virome epidemiology and systematic analysis of virus-host interactions.

bioinformatics↗

Assessing computational predictions of antimicrobial resistance phenotypes from microbial genomes

The advent of rapid whole-genome sequencing has created new opportunities for computational prediction of antimicrobial resistance (AMR) phenotypes from genomic data. Both rule-based and machine learning (ML) approaches have been explored for this task, but systematic benchmarking is still needed. Here, we evaluated four state-of-the-art ML methods (Kover, PhenotypeSeeker, Seq2Geno2Pheno, and Aytan-Aktug), an ML baseline, and the rule-based ResFinder by training and testing each of them across 78 species-antibiotic datasets, using a rigorous benchmarking workflow that integrates three evaluation approaches, each paired with three distinct sample splitting methods. Our analysis revealed considerable variation in the performance across techniques and datasets. Whereas ML methods generally excelled for closely related strains, ResFinder excelled for handling divergent genomes. Overall, Kover most frequently ranked top among the ML approaches, followed by PhenotypeSeeker and Seq2Geno2Pheno. AMR phenotypes for antibiotic classes such as macrolides and sulfonamides were predicted with the highest accuracies. The quality of predictions varied substantially across species-antibiotic combinations, particularly for beta-lactams; across species, resistance phenotyping of the beta-lactams compound, aztreonam, amox-clav, cefoxitin, ceftazidime, and piperacillin/tazobactam, alongside tetracyclines demonstrated more variable performance than the other benchmarked antibiotics. By organism, C. jejuni and E. faecium phenotypes were more robustly predicted than those of Escherichia coli, Staphylococcus aureus, Salmonella enterica, Neisseria gonorrhoeae, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Streptococcus pneumoniae, and Mycobacterium tuberculosis. In addition, our study provides software recommendations for each species-antibiotic combination. It furthermore highlights the need for optimization for robust clinical applications, particularly for strains that diverge substantially from those used for training.

bioinformatics↗

Nucleotide-pair encoding of 16S rRNA sequences for host phenotype and biomarker detection

Identifying combinations of taxa distinctive for microbiome-associated diseases is considered key to the establishment of diagnosis and therapy options in precision medicine and imposes high demands on accuracy of microbiome analysis techniques. We propose subsequence based 16S rRNA data analysis, as a new paradigm for microbiome phenotype classification and biomarker detection. This method and software called DiTaxa substitutes standard OTU-clustering or sequence-level analysis by segmenting 16S rRNA reads into the most frequent variable-length subsequences. These subsequences are then used as data representation for downstream phenotype prediction, biomarker detection and taxonomic analysis. Our proposed sequence segmentation called nucleotide-pair encoding (NPE) is an unsupervised data-driven segmentation inspired by Byte-pair encoding, a data compression algorithm. The identified subsequences represent commonly occurring sequence portions, which we found to be distinctive for taxa at varying evolutionary distances and highly informative for predicting host phenotypes. We compared the performance of DiTaxa to the state-of-the-art methods in disease phenotype prediction and biomarker detection, using human-associated 16S rRNA samples for periodontal disease, rheumatoid arthritis and inflammatory bowel diseases, as well as a synthetic benchmark dataset. DiTaxa identified 17 out of 29 taxa with confirmed links to periodontitis (recall= 0.59), relative to 3 out of 29 taxa (recall= 0.10) by the state-of-the-art method. On synthetic benchmark data, DiTaxa obtained full precision and recall in biomarker detection, compared to 0.91 and 0.90, respectively. In addition, machine-learning classifiers trained to predict host disease phenotypes based on the NPE representation performed competitively to the state-of-the art using OTUs or k-mers. For the rheumatoid arthritis dataset, DiTaxa substantially outperformed OTU features with a macro-F1 score of 0.76 compared to 0.65. Due to the alignment- and reference free nature, DiTaxa can efficiently run on large datasets. The full analysis of a large 16S rRNA dataset of 1359 samples required {approx}1.5 hours on 20 cores, while the standard pipeline needed {approx}6.5 hours in the same setting.\n\nAvailabilityAn implementation of our method called DiTaxa is available under the Apache 2 licence at http://llp.berkeley.edu/ditaxa.

bioinformatics↗