Search bioRxivSearch

Biology subjects

Deogun, J. S.

Publications and source records attributed to Deogun, J. S..

2 recordsLinked to original sources

A consensus-based ensemble approach to improve de novo transcriptome assembly

BackgroundSystems-level analyses, such as differential gene expression analysis, co-expression analysis, and metabolic pathway reconstruction, depend on the accuracy of the transcriptome. Multiple tools exist to perform transcriptome assembly from RNAseq data. However, assembling high quality transcriptomes is still not a trivial problem. This is especially the case for non-model organisms where adequate reference genomes are often not available. Different methods produce different transcriptome models and there is no easy way to determine which are more accurate. Furthermore, having alternative-splicing events exacerbates such difficult assembly problems. While benchmarking transcriptome assemblies is critical, this is also not trivial due to the general lack of true reference transcriptomes. ResultsIn this study, we first provide a pipeline to generate a set of the benchmark transcriptome and corresponding RNAseq data. Using the simulated benchmarking datasets, we compared the performance of various transcriptome assembly approaches including both de novo and genome-guided methods. The results showed that the assembly performance deteriorates significantly when alternative transcripts (isoforms) exist or for genome-guided methods when the reference is not available from the same genome. To improve the transcriptome assembly performance, leveraging the overlapping predictions between different assemblies, we present a new consensus-based ensemble transcriptome assembly approach, ConSemble. ConclusionsWithout using a reference genome, ConSemble using four de novo assemblers achieved an accuracy up to twice as high as any de novo assemblers we compared. When a reference genome is available, ConSemble using four genome-guided assemblies removed many incorrectly assembled contigs with minimal impact on correctly assembled contigs, achieving higher precision and accuracy than individual genome-guided methods. Furthermore, ConSemble using de novo assemblers matched or exceeded the best performing genome-guided assemblers even when the transcriptomes included isoforms. We thus demonstrated that the ConSemble consensus strategy both for de novo and genome-guided assemblers can improve transcriptome assembly. The RNAseq simulation pipeline, the benchmark transcriptome datasets, and the script to perform the ConSemble assembly are all freely available from: http://bioinfolab.unl.edu/emlab/consemble/.

bioinformatics

STAG-CNS: An Order-Aware Conserved Non-coding Sequences Discovery Tool For Arbitrary Numbers of Species

One method for identifying noncoding regulatory regions of a genome is to quantify rates of divergence between related species, as functional sequence will generally diverge more slowly. Most approaches to identifying these conserved noncoding sequences (CNS) based on alignment have had relatively large minimum sequence lengths ([>=]15 base pair) compared to the average length of known transcription factor binding sites. To circumvent this constraint, STAG-CNS integrates data from the promoters of conserved orthologous genes in three or more species simultaneously. Using data from up to six grass species made it possible to identify conserved sequences as short at 9 base pairs with FDP [<=] 0.05. These CNS exhibit greater overlap with open chromatin regions identified using DNase I hypersensitivity, and are enriched in the promoters of genes involved in transcriptional regulation. STAG-CNS was further employed to characterize loss of conserved noncoding sequences associated with retained duplicate genes from the ancient maize polyploidy. Genes with fewer retained CNS show lower overall expression, although this bias is more apparent in samples of complex organ systems containing many cell types, suggesting CNS loss may correspond to a reduced number of expression contexts rather than lower expression levels across the entire ancestral expression domain.

bioinformatics