Search bioRxivSearch

Biology subjects

Sabot, F.

Publications and source records attributed to Sabot, F..

3 recordsLinked to original sources

TOGGLe, a flexible framework for easily building complex workflows and performing robust large-scale NGS analyses

The advent of NGS has intensified the need for robust pipelines to perform high-performance automated analyses. The required softwares depend on the sequencing method used to produce raw data (e.g. Whole genome sequencing, Genotyping By Sequencing, RNASeq) as well as the kind of analyses to carry on (GWAS, population structure, differential expression). These tools have to be generic and scalable, and should meet the biologists needs.\n\nHere, we present the new version of TOGGLe (Toolbox for Generic NGS Analyses), a simple and highly flexible framework to easily and quickly generate pipelines for large-scale second-and third-generation sequencing analyses, including multi-threading support. TOGGLe comprises a workflow manager designed to be as effortless as possible to use for biologists, so the focus can remain on the analyses. Embedded pipelines are easily customizable and supported analyses are reproducible and shareable. TOGGLe is designed as a generic, adaptable and fast evolutive solution, and has been tested and used in large-scale projects with numerous samples and organisms. It is freely available at http://toggle.southgreen.fr/ under the GNU GPLv3/CeCill-C licenses) and can be deployed onto HPC clusters as well as on local machines.

bioinformatics

Comparison of two African rice species through a new pan-genomic approach on massive data

Pangenome theory implies that individuals from a given group/species share only a given part of their genome (core-genome), the remaining part being the dispensable one. Domestication process implies a small number of founder individuals, and thus a large core-genome compared to dispensable at the first steps of domestication. We sequenced at high depth 120 cultivated African rice Oryza glaberrima and of 74 wild relatives O. barthii, and mapped them on the external reference from Asian rice O. sativa. We then use a novel DepthOfCoverage approach to identif missing genes. After comparing the two species, we shown that the cultivated species has a smaller core-genome than the wild one, as well as an expected smaller dispensable one. This unexpected output however replaces in perspective the inadequacy of cultivated crops to wilderness.

genomics

Evolutionary forces affecting synonymous variations in plant genomes

Base composition is highly variable among and within plant genomes, especially at third codon positions, ranging from GC-poor and homogeneous species to GC-rich and highly heterogeneous ones (particularly Monocots). Consequently, synonymous codon usage is biased in most species, even when base composition is relatively homogeneous. The causes of these variations are still under debate, with three main forces being possibly involved: mutational bias, selection and GC-biased gene conversion (gBGC). So far, both selection and gBGC have been detected in some species but how their relative strength varies among and within species remains unclear. Population genetics approaches allow to jointly estimating the intensity of selection, gBGC and mutational bias. We extended a recently developed method and applied it to a large population genomic datasets based on transcriptome sequencing of 11 angiosperm species spread across the phylogeny. We found that base composition is far from mutation-drift equilibrium in most genomes and that gBGC is a widespread and stronger process than selection. gBGC could strongly contribute to base composition variation among plant species, implying that it should be taken into account in plant genome analyses, especially for GC-rich ones.

evolutionary biology