Search bioRxivSearch

Biology subjects

Lemaitre, C.

Publications and source records attributed to Lemaitre, C..

4 recordsLinked to original sources

DiscoSnp-RAD: de novo detection of small variants for population genomics

We present an original method to de novo call variants for Restriction site associated DNA Sequencing (RAD-Seq). RAD-Seq is a technique characterized by the sequencing of specific loci along the genome, that is widely employed in the field of evolutionary biology since it allows to exploit variants (mainly SNPs) information from entire populations at a reduced cost. Common RAD dedicated tools, as STACKS or IPyRAD, are based on all-versus-all read comparisons, which require consequent time and computing resources. Based on the variant caller DiscoSnp, initially designed for shotgun sequencing, DiscoSnp-RAD avoids this pitfall as variants are detected by exploring the De Bruijn Graph built from all the read datasets. We tested the implementation on RAD data from 259 specimens of Chiastocheta flies, morphologically assigned to 7 species. All individuals were successfully assigned to their species using both STRUCTURE and Maximum Likelihood phylogenetic reconstruction. Moreover, identified variants succeeded to reveal a within species structuration and the existence of two populations linked to their geographic distributions. Furthermore, our results show that DiscoSnp-RAD is at least one order of magnitude faster than state-of-the-art tools. The overall results show that DiscoSnp-RAD is suitable to identify variants from RAD data, and stands out from other tools due to his completely different principle, making it significantly faster, in particular on large datasets.\n\nLicenseGNU Affero general public license\n\nAvailabilityhttps://github.com/GATB/DiscoSnp\n\nContactjeremy.gauthier@inria.fr

bioinformatics

DiscoSnp++: de novo detection of small variants from raw unassembled read set(s)

MotivationNext Generation Sequencing (NGS) data provide an unprecedented access to life mechanisms. In particular, these data enable to detect polymorphisms such as SNPs and indels. As these polymorphisms represent a fundamental source of information in agronomy, environment or medicine, their detection in NGS data is now a routine task. The main methods for their prediction usually need a reference genome. However, non-model organisms and highly divergent genomes such as in cancer studies are extensively investigated.\n\nResultsWe propose DiscoSnp++, in which we revisit the DiscoSnp algorithm. DiscoSnp++ is designed for detecting and ranking all kinds of SNPs and small indels from raw read set(s). It outputs files in fasta and VCF formats. In particular, predicted variants can be automatically localized afterwards on a reference genome if available. Its usage is extremely simple and its low resource requirements make it usable on common desktop computers. Results show that DiscoSnp++ performs better than state-of-the-art methods in terms of computational resources and in terms of results quality. An important novelty is the de novo detection of indels, for which we obtained 99% precision when calling indels on simulated human datasets and 90% recall on high confident indels from the Platinum dataset.\n\nLicenseGNU Affero general public license\n\nAvailabilityhttps://github.com/GATB/DiscoSnp\n\nContactpierre.peterlongo@inria.fr

bioinformatics

Disentangling The Causes For Faster-X Evolution In Aphids

Faster evolution of X chromosomes has been documented in several species and results from the increased efficiency of selection on recessive alleles in hemizygous males and/or from increased drift due to the smaller effective population size of X chromosomes. Aphids are excellent models for evaluating the importance of selection in faster-X evolution, because their peculiar life-cycle and unusual inheritance of sex-chromosomes lead to equal effective population sizes for X and autosomes. Because we lack a high-density genetic map for the pea aphid whose complete genome has been sequenced, we assigned its entire genome to the X and autosomes based on ratios of sequencing depth in males and females. Unexpectedly, we found frequent scaffold misassembly, but we could unambiguously locate 13,726 genes on the X and 19,263 on autosomes. We found higher non-synonymous to synonymous substitutions ratios (dN/dS) for X-linked than for autosomal genes. Our analyses of substitution rates together with polymorphism and expression data showed that relaxed selection is likely to contribute predominantly to faster-X as a large fraction of X-linked genes are expressed at low rates and thus escape selection. Yet, a minor role for positive selection is also suggested by the difference between substitution rates for X and autosomes for male-biased genes (but not for asexual female-biased genes) and by lower Tajimas D for X-linked than for autosomal genes with highly male-biased expression patterns. This study highlights the relevance of organisms displaying alternative inheritance of chromosomes to the understanding of forces shaping genome evolution.

evolutionary biology

Critical Assessment of Metagenome Interpretation - a benchmark of computational metagenomics software

In metagenome analysis, computational methods for assembly, taxonomic profiling and binning are key components facilitating downstream biological data interpretation. However, a lack of consensus about benchmarking datasets and evaluation metrics complicates proper performance assessment. The Critical Assessment of Metagenome Interpretation (CAMI) challenge has engaged the global developer community to benchmark their programs on datasets of unprecedented complexity and realism. Benchmark metagenomes were generated from ~700 newly sequenced microorganisms and ~600 novel viruses and plasmids, including genomes with varying degrees of relatedness to each other and to publicly available ones and representing common experimental setups. Across all datasets, assembly and genome binning programs performed well for species represented by individual genomes, while performance was substantially affected by the presence of related strains. Taxonomic profiling and binning programs were proficient at high taxonomic ranks, with a notable performance decrease below the family level. Parameter settings substantially impacted performances, underscoring the importance of program reproducibility. While highlighting current challenges in computational metagenomics, the CAMI results provide a roadmap for software selection to answer specific research questions.

bioinformatics