Search bioRxiv⌕ Search

Biology subjects

Rozwalak, P.

Publications and source records attributed to Rozwalak, P..

4 recordsLinked to original sources

AncientMetagenomeDir dating metadataset highlights need for standardised radiocarbon reporting in ancient DNA

Ancient DNA is a valuable data source for the understanding of our past. However, to effectively interpret this data, it is essential to know the age of the samples from which the DNA is obtained. Although the field of palaeogenomics has been recognised for its robust open data sharing practices, dating information associated with analysed samples is not reported consistently across palaeogenomic studies, nor is it included as metadata in most genetic data repositories. Here, we describe the addition of standardised precise dating information for ancient microbial genomes into the AncientMetagenomeDir metadata repository of published ancient metagenomic samples. This extension currently includes dating information for over 700 ancient microbial genomic datasets, of which 333 are dated using historical, contextual, or stratigraphic methods, and 405 are radiocarbon dated. We quantitatively assess the quality of radiocarbon date reporting and find that, despite established reporting conventions, radiocarbon dating information is often reported inconsistently across ancient metagenomic studies. This new resource provides ancient microbial researchers with standardised dating information that facilitates more accurate and consistent analysis of metagenomic sequencing data. The dataset also highlights the need for greater standardisation of radiocarbon date reporting in original publications in order to allow effective reuse of this and future ancient microbial data.

bioinformatics↗

Jaeger: an accurate and fast deep-learning tool to detect bacteriophage sequences

Viruses are integral to every biome on Earth, yet we still need a more comprehensive picture of their identity and global distribution. Global metagenomics sequencing efforts revealed the genomic content of tens of thousands of environmental samples, however identifying the viral sequences in these datasets remains challenging due to their vast genomic diversity. Here, we address identifying bacteriophage sequences in unlabeled sequencing data. In a recent benchmarking paper, we observed that existing deep-learning tools show a high true positive rate, but may also produce many false positives when confronted with divergent sequences. To tackle this challenge, we introduce Jaeger, a novel deep-learning method designed specifically for identifying bacteriophage genome fragments. Extensive benchmarking on the IMG/VR database and real-world metagenomes reveals Jaegers consistent high sensitivity (0.87) and precision (0.92). Applying Jaeger to over 16,000 metagenomic assemblies from the MGnify database yielded over five million putative phage contigs. On average, Jaeger is around 20 times faster than the other state-of-the-art methods. Jaeger is available at https://github.com/MGXlab/Jaeger.

bioinformatics↗

Ultrafast and accurate sequence alignment and clustering of viral genomes

Viromics produces millions of viral genomes and fragments annually, overwhelming traditional sequence comparison methods. We introduce Vclust, a novel approach that determines average nucleotide identity by Lempel-Ziv parsing and clusters viral genomes with thresholds endorsed by authoritative viral genomics and taxonomy consortia. Vclust demonstrates superior accuracy and efficiency compared to existing tools, clustering millions of virus genomes in a few hours on a mid-range workstation.

bioinformatics↗

Ultra-conserved bacteriophage genome sequence identified in 1300-year-old human paleofaeces

Bacteriophages are widely recognised as rapidly evolving biological entities. However, we discovered an ancient genome nearly identical to present-day Mushuvirus mushu, a phage that infects commensal microorganisms in the human gut ecosystem. The DNA damage patterns of this genome have confirmed its ancient origin, and, despite 1300 years of evolution, the ancient Mushuvirus genome shares 97.7% nucleotide identity with its modern counterpart, indicating a long-term relationship between the prophage and its host. We also reconstructed and authenticated 297 other phage genomes from the last 5300 years, including those belonging to unknown families. Our findings demonstrate the feasibility of reconstructing ancient phage genomes, expanding the known virosphere, and offering new insights into phage-bacteria interactions that cover several millennia.

microbiology↗