Search bioRxivSearch

Biology subjects

Dereeper, A.

Publications and source records attributed to Dereeper, A..

2 recordsLinked to original sources

Rice Galaxy: an open resource for plant science

Background\n\nRice molecular genetics, breeding, genetic diversity, and allied research (such as rice-pathogen interaction) have adopted sequencing technologies and high density genotyping platforms for genome variation analysis and gene discovery. Germplasm collections representing rice diversity, improved varieties and elite breeding materials are accessible through rice gene banks for use in research and breeding, with many having genome sequences and high density genotype data available. Combining phenotypic and genotypic information on these accessions enables genome-wide association analysis, which is driving quantitative trait loci (QTL) discovery and molecular marker development. Comparative sequence analyses across QTL regions facilitate the discovery of novel alleles. Analyses involving DNA sequences and large genotyping matrices for thousands of samples, however, pose a challenge to non-computer savvy rice researchers.\n\nFindings\n\nWe adopted the Galaxy framework to build the federated Rice Galaxy resource, with shared datasets, tools, and analysis workflows relevant to rice research. The shared datasets include high density genotypes from the 3,000 Rice Genomes project and sequences with corresponding annotations from nine published rice genomes. Rice Galaxy includes tools for designing single nucleotide polymorphism (SNP) assays, analyzing genome-wide association studies, population diversity, rice-bacterial pathogen diagnostics, and a suite of published genomic prediction methods. A prototype Rice Galaxy compliant to Open Access, Open Data, and Findable, Accessible, Interoperable, and Reproducible principles is also presented.\n\nConclusions\n\nRice Galaxy is a freely available resource that empowers the plant research community to perform state-of-the-art analyses and utilize publicly available big datasets for both fundamental and applied science.

bioinformatics

TOGGLe, a flexible framework for easily building complex workflows and performing robust large-scale NGS analyses

The advent of NGS has intensified the need for robust pipelines to perform high-performance automated analyses. The required softwares depend on the sequencing method used to produce raw data (e.g. Whole genome sequencing, Genotyping By Sequencing, RNASeq) as well as the kind of analyses to carry on (GWAS, population structure, differential expression). These tools have to be generic and scalable, and should meet the biologists needs.\n\nHere, we present the new version of TOGGLe (Toolbox for Generic NGS Analyses), a simple and highly flexible framework to easily and quickly generate pipelines for large-scale second-and third-generation sequencing analyses, including multi-threading support. TOGGLe comprises a workflow manager designed to be as effortless as possible to use for biologists, so the focus can remain on the analyses. Embedded pipelines are easily customizable and supported analyses are reproducible and shareable. TOGGLe is designed as a generic, adaptable and fast evolutive solution, and has been tested and used in large-scale projects with numerous samples and organisms. It is freely available at http://toggle.southgreen.fr/ under the GNU GPLv3/CeCill-C licenses) and can be deployed onto HPC clusters as well as on local machines.

bioinformatics