Search bioRxiv⌕ Search

Biology subjects

Raffestin, B.

Publications and source records attributed to Raffestin, B..

2 recordsLinked to original sources

Logan: Planetary-Scale Genome Assembly Surveys Life's Diversity

The breadth of lifes diversity is unfathomable, but public nucleic acid sequencing data offers a window into the dispersion and evolution of genetic diversity across Earth. However the rapid growth and accumulation of sequence data have outpaced efficient analysis capabilities. The largest collection of freely available sequencing data is the Sequence Read Archive (SRA), comprising 27.3 million datasets or 5 x 1016 basepairs. To realize the potential of the SRA, we constructed Logan, a massive sequence assembly transforming short reads into long contigs and compressing the data over 100-fold, enabling highly efficient petabase-scale analysis. We created Logan-Search, a k-mer index of Logan for free planetary-scale sequence search, returning matches in minutes. We used Logan contigs to identify >200 million plastic-degrading enzyme homologs, and validate novel enzymes with catalytic activities exceeding current reference standards. Further, we vastly expand the known diversity of proteins (30-fold over UniRef50), plasmids (22-fold over PLSDB), P4 satellites (4.5-fold), and the recently described Obelisk RNA elements (3.7-fold). Logan also enables ecological and biomedical data mining, such as global tracking of antimicrobial resistance genes and the characterization of viral reactivation across millions of human BioSamples. By transforming the SRA, Logan democratizes access to the worlds public genetic data and opens frontiers in biotechnology, molecular ecology, and global health.

bioinformatics↗

Bacterial strain nomenclature in the genomic era: Life Identification Numbers using a gene-by-gene approach

Unified strain taxonomies are needed for the epidemiological surveillance of bacterial pathogens and international communication in microbiological research. Core genome multilocus sequence typing (cgMLST) holds great promise for standardized high-resolution strain genotyping. However, this approach faces challenges including classification instability and disconnection of new nomenclature from widely adopted classical MLST identifiers. This essay discusses the cgMLST-based Life Identification Number (LIN) method, recently proposed as a stable multilevel strain taxonomy system applicable to most bacterial pathogens. We describe how LIN codes are implemented and used in practice for precise strain definitions and epidemiological tracking. Glossary Multilocus sequence typing (MLST)A genotyping method applied mostly to microbial strains to study population structure and epidemiology, based on comparing the nucleotide sequences of a small number (typically seven) of housekeeping protein-coding genes. In MLST, allele numbers are assigned to each sequence variant (allele) of a given gene. The MLST genotype of a bacterial strain is defined by the combination of the allele numbers observed at the genes that are included in the genotyping scheme. A sequence type (ST) is assigned to each unique combination of alleles, called an MLST profile. MLST was invented in 1998 and became a de-facto standard taxonomy of bacterial strains, albeit at low resolution. Core genome MLSTAn extension of MLST that analyzes sequence variation across hundreds to thousands of conserved (core) genes, shared by all strains of a species, providing higher resolution typing for genomic epidemiology and evolutionary studies. cgMLST schemes typically comprise 2000 to 4000 genes, depending on the genome size and genetic variation (in terms of presence/absence of genes) within bacterial species. A core genome sequence type (cgST) can be assigned to unique cgMLST profiles, i.e., a unique combination of cgMLST allelic numbers. Whole Genome Sequencing (WGS)A method that determines the complete DNA sequence of an organisms genome in a single process, providing comprehensive information for comparative genetic analyses based on cgMLST or other analytic methods. Single Nucleotide Polymorphisms (SNPs)Variations at a single base position in the DNA sequence among individuals isolates, strains or species, used as genetic markers for studying for example, evolutionary relationships or strain identity. Average nucleotide identity (ANI)A measure of genomic similarity between two organisms, calculated as the average percentage of identical nucleotides in orthologous genomic regions; commonly used to assess species-level relatedness in prokaryotes. TaxonomyHere, we apply the word taxonomy to bacterial strains as a system of classifying, naming and identifying strains based on shared genetic characteristics as defined by e.g., cgMLST.

microbiology↗