Search bioRxiv⌕ Search

Biology subjects

Jaques, G.

Publications and source records attributed to Jaques, G..

2 recordsLinked to original sources

GenomeCompendium: A database for the integrated analysis of repeats, assembly quality and functional content of complete prokaryotic genomes

Microorganisms hold great promise for urgent global needs such as increasing sustainable agricultural production while reducing chemical fertilizer and pesticide use or providing novel classes of antimicrobials/therapeutics. Moving from analyzing microbiome composition to applying synthetic communities and studying their functions requires access to isolates and complete genome sequences. By spanning the frequent repeats, long-read sequencing can resolve complex prokaryotic genomes, yet error-prone short-read assemblies dominate. We here release the GenomeCompendium, a public database and interactive analysis tool for complete prokaryotic genomes (https://genome-compendium.com/). Using NCBI RefSeq (~47,000) and GenBank (~13,000) genomes, we integrated available metadata, GTDB taxonomy and computed features including repeat classes and gene content screening, intragenomic 16S rRNA sequence identity, and biosynthetic gene cluster co-occurrences. Evaluating repeat content and assembly complexity metrics, we identify taxonomic ranks dominated by difficult-to-assemble genomes and show that complex, repeat-rich genomes are more common than previously estimated. By mining metadata, our quality control flags 6.3% of RefSeq assemblies as potentially erroneous or incomplete. As valuable reference for data mining and to track taxonomic coverage, the GenomeCompendium links ~90 features across genomes, offers downloadable reports and -as unique features- pre-computed proteogenomics databases to improve genome annotations of RefSeq strains and the ability to analyze any uploaded prokaryotic genome.

genomics↗

A revised genome annotation of the model cyanobacterium Synechocystis based on start and stop codon-enriched ribosome profiling and proteogenomics

Cyanobacteria are important primary producers and are used as microbial cell factories due to their ability to use solar light for oxygenic photosynthesis. Synechocystis sp. PCC 6803 is a popular model cyanobacterium, yet there are ambiguities in the precise coding regions of many genes, and numerous genes encoding small proteins have remained undetected. Here we present the results of a riboproteogenomic analysis, combining ribosome profiling (Ribo-seq) analysis involving inhibitors that stall ribosomes at translation initiation and termination sites (TIS- and TTS-Ribo-seq), with a proteogenomic reevaluation and reannotation of its entire genome. We report evidence for the translation of 3,055 annotated genes based on proteogenomics (83%), of 3,492 based on Ribo-seq (95.2%), and of 3,018 supported by both methods (82%). The data suggested unannotated protein-coding genes and corrections for annotated ones. We validated 15 small proteins translated from antisense RNAs, from intergenic and intragenic regions and provide proteogenomic support for up to 69 further, mostly small proteins. With slr0489, slr1079 and slr1082 we identified three genes with intragenic out-of-frame translons and show that both the internal and host reading frames are translated and that the resulting proteins interact with each other. Our data can be accessed via an intuitive and interactive genome browser platform at https://www.bioinf.uni-freiburg.de/[~]ribobase/. They illustrate the enormous value of consolidating genome annotations in the context of integrated experimental data.

microbiology↗