Search bioRxivSearch

Biology subjects

Martin, J.-F.

Publications and source records attributed to Martin, J.-F..

2 recordsLinked to original sources

High throughput amplicon sequencing to assess within- and between-host genetic diversity in plant viruses

Molecular epidemiology approaches at the landscape scale require to study the genetic diversity of viral populations from numerous hosts and to characterize mixed infections. In such a context, high-throughput amplicon sequencing (HTAS) techniques create interesting opportunities as they allow identifying distinct variants within a same host while simultaneously genotyping a high number of samples. Validating variants produced by HTAS may, however, remain difficult due to biases occurring at different steps of the data-generating process (e.g. environmental contaminations and sequencing error). Here, we focused on Endive necrotic mosaic virus (ENMV), a member of family Potyviridae, genus Potyvirus to develop an HTAS approach and to characterize the genetic diversity at the intra- and inter-host levels from 430 samples collected over an area of 1660 km2 located in south-eastern France. We demonstrated how it is possible, by incorporating various controls in the experimental design and by performing independent sample replicates, to estimate potential biases in HTAS results and to implement an automated and robust variant calling procedure.\n\nHighlightsO_LIHigh-throughput amplicon sequencing to assess plant virus genetic diversity\nC_LIO_LIEstimating bias in high throughput amplicon sequencing results\nC_LIO_LIAutomated variant calling procedure for robust high throughput amplicon sequencing\nC_LI

molecular biology

Challenges and solutions for transcriptome assembly in non-model organisms with an application to hybrid specimens

Analyses of high-throughput transcriptome sequences of non-model organisms are based on two main approaches: de novo assembly and genome-guided assembly using mapping to assign reads prior to assembly. Given the limits of mapping reads to a reference when it is highly divergent, as is frequently the case for non-model species, we evaluate whether using blastn would outperform mapping methods for read assignment in such situations (>15% divergence). We demonstrate its high performance by using simulated reads of lengths corresponding to those generated by the most common sequencing platforms, and over a realistic range of genetic divergence (0% to 30% divergence). Here we focus on gene identification and not on resolving the whole set of transcripts (i.e. the complete transcriptome). For simulated datasets, the transcriptome-guided assembly based on blastn recovers 94.8% of genes irrespective of read length at 0% divergence; however, assignment rate of reads is negatively correlated with both increasing divergence level and reducing read lengths. Nevertheless, we still observe 92.6% of recovered genes at 30% divergence irrespective of read length. This analysis also produces a categorization of genes relative to their assignment, and suggests guidelines for data processing prior to analyses of comparative transcriptomics and gene expression to minimize potential inferential bias associated with incorrect transcript assignment. We also compare the performances of de novo assembly alone vs in combination with a transcriptome-guided assembly based on blastn via simulation and empirically, using data from a cyprinid fish species and from an oak species. For any simulated scenario, the transcriptome-guided assembly using blastn outperforms the de novo approach alone, including when the divergence level is beyond the reach of mapping methods. Combining de novo assembly and a related reference transcriptome for read assignment also addresses the bias/error in contigs caused by the dependence on a related reference alone. Empirical data corroborate those findings when assembling transcriptomes from the two non-model organisms: Parachondrostoma toxostoma (fish) and Quercus pubescens (plant). For the fish species, out of the 31,944 genes known from D. rerio, the guided and de novo assemblies recover respectively 20,605 and 20,032 genes but the performance of the guided assembly approach is much higher for both the contiguity and completeness metrics. For the oak, out of the 29,971 genes known from Vitis vinifera, the transcriptome-guided and de novo assemblies display similar performance but the new guided approach detects 16,326 genes where the de novo assembly only detects 9,385 genes.

genomics