Search bioRxiv⌕ Search

Biology subjects

Senelle, G.

Publications and source records attributed to Senelle, G..

2 recordsLinked to original sources

TB-annotator: a scalable web application that allows in-depth analysis of very large sets of publicly available Mycobacterium tuberculosis complex genomes

Tuberculosis continues to be one of the most threatening bacterial diseases in the world. However, we currently have more than 160,000 Short Read Archives (SRAs) of Mycobacterium tuberculosis complex. Such a large amount of data should help to the understanding and the fight against this bacterium. To accomplish this, it would be necessary to thoroughly and comprehensively examine this significant mass of data. This is what TB-Annotator proposes to do, combining a database containing all the diversity of these 160,000 SRAs (at least, SRAs with a reasonable read size and quality), and a fully featured analysis platform to explore and query such a large amount of data. The objective of this article is to present this platform centered on the key notion of exclusivity, to show its numerous capacities (detection of single nucleotide variants, insertion sequences, deletion regions, spoligotyping, etc.) and its general functioning. We will compare TB-Annotator to existing tools for the study of tuberculosis, and show that its objectives are original and have no equivalent at present. The database on which it is based will be presented, with the numerous advanced search queries and screening capacities it offers, and the interest and originality of its phylogenetic tree navigation interface will be detailed. We will end this article with examples of the achievements made possible by the TB-Annotator, followed by avenues for future improvement.

microbiology↗

An updated evolutionary history and taxonomy of Mycobacterium tuberculosis lineage 5, also called M. africanum

Contrarily to other lineages such as L2 and L4, there are still scarce whole-genome-sequence data on L5-L6 MTBC clinical isolates in public genomes repositories. Recent results suggest a high complexity of L5 history in Africa. It is of importance for an adequate assessment of TB infection in Africa, that is still related to the presence of L5-L6 MTBC strains. This study reports a significant improvement of our knowledge of L5 diversity, phylogeographical history, and global population structure of Mycobacterium africanum L5. To achieve this aim, we sequenced new clinical isolates from Northern Nigeria and from proprietary collections, and used a new powerful bioinformatical pipeline, TB-Annotator that explores not only the shared SNPs but also shared missing genes, identical IS6110 insertion sites and shared regions of deletion. This study using both newly sequenced genomes and available public genomes allows to describe new L5 sublineages. We report that the MTBC L5 tree is made-up of at least 12 sublineages from which 6 are new descriptions. We confront our new classification to the most recent published one and suggest new naming for the discovered sublineages. Finally, we discuss the phylogeographical specificity of sublineages 5.1 and sublineage 5.2 and suggest a new hypothesis of L5-L6 emergence in Africa. Impact statementRecent studies on Mycobacterium africanum (L5-L6-L9 of MTBC) genomic diversity and its evolution in Africa discovered three new lineages of the Mycobacterium tuberculosis complex (MTBC) in the last ten years (L7, L8, L9). These discoveries are symptomatic of the delay in characterizing the diversity of the MTBC on the African continent. Another understudied part of MTBC diversity is the intra-lineage diversity of L5 and L6. This study unravels an hidden diversity of L5 in Africa and provides a more exhaustive description of specific genetic features of each sublineage by using a proprietary "TB-Annotator" pipeline. Furthermore, we identify different phylogeographical localization trends between L5.1 and L5.2, suggesting different histories. Our results suggest that a better understanding of the spatiotemporal dynamics of MTBC in Africa absolutely requires a large sampling effort and powerful tools to dig into the retrieved diversity. Data summary[A section describing all supporting external data, software or code, including the DOI(s) and/or accession numbers(s), and the associated URL. If no data was generated or reused in the research, please state this.] The search was done in the TB-Annotator 15901 genomes version which is available at: http://(to be added). The new sequenced genomes are available via NCBI Bioproject accession number: (to be added). The authors confirm all supporting data, code and protocols have been provided within the article or through supplementary data files.

genomics↗