Search bioRxiv⌕ Search

Biology subjects

Brasseur, M. V.

Publications and source records attributed to Brasseur, M. V..

2 recordsLinked to original sources

Comparative Analysis of De Novo Assemblers and Quantification Software for RNA-sequencing Data in Non-Model Arthropods

BackgroundRNA-sequencing has greatly improved our understanding of the transcriptomic regulation of fundamental biological processes. Although the method has matured significantly within the last decade, bioinformatic processing of the resulting high-dimensional data sets is still challenging and the performance of algorithms can vary between data sets. As a consequence, for most non-model organisms, in particular arthropods, there is no or limited literature evidence which software is best suited to handle taxon-specific data characteristics. Therefore, we evaluated the performance of different de nonvo transcriptome assembler (Trinity, rnaSPAdes, IDBA-tran) and transcript quantification software (RSEM, Salmon) on transcriptomic data of a non-model insect and freshwater crustacean species, as well as the impact of different quality trimming strategies on the downstream bioinformatic processing results. ResultsWhile the trimming strategy had no considerable effect on the quality of transcriptome assemblies, the choice of the assembler had a substantial impact. IDBA-tran was less sensitive than the two other assemblers and produced the most fragmented transcriptome assemblies. The low remapping rates of reads against IDBA-tran assemblies further suggest that the input read data was not effectively leveraged by this algorithm. In contrast, Trinity and rnaSPAdes both generated comprehensive and contiguous de novo transcriptome assemblies, although Trinity appeared to be slightly more sensitive. This increased sensitivity, however, was associated with a higher redundancy in Trinity-generated assemblies compared to assemblies produced with rnaSPAdes. When the quality of the transcriptome assembly was high, RSEM and Salmon were able to identify the origin of at least 90% of the read data in the reference. Despite their different underlying quantification approaches, the estimated transcript counts of both tools were highly correlated and their expression signal was consistent. Notably, the alignment-free quantification algorithm Salmon was substantially faster than the alignment-based approach of RSEM. Furthermore, it was also slightly more sensitive, increasing the average re-mapping rate to [~]98%. ConclusionSince the performance of bioinformatic algorithms, especially of de novo assemblers, varies for different RNA-sequencing data sets, establishing an appropriate analysis workflow remains an important task. Our results show that the better performing combinations of algorithms produce congruent count data sets with consistent expression signal, highlighting the robustness of RNA-sequencing data analysis software.

bioinformatics↗

Predicting environmental stressor levels with machine learning: a comparison between amplicon sequencing, metagenomics, and total RNA sequencing based on taxonomically assigned data

BackgroundMicrobes are increasingly (re)considered for environmental assessments because they are powerful indicators for the health of ecosystems. The complexity of microbial communities necessitates powerful novel tools to derive conclusions for environmental decision-makers, and machine learning is a promising option in that context. While amplicon sequencing is typically applied to assess microbial communities, metagenomics and total RNA sequencing (herein summarized as omics-based methods) can provide a more holistic picture of microbial biodiversity at sufficient sequencing depths. Despite this advantage, amplicon sequencing and omics-based methods have not yet been compared for taxonomy-based environmental assessments with machine learning. In this study, we applied 16S and ITS-2 sequencing, metagenomics, and total RNA sequencing to samples from a stream mesocosm experiment that investigated the impacts of two aquatic stressors, insecticide and increased fine sediment deposition, on stream biodiversity. We processed the data using similarity clustering and denoising (only applicable to amplicon sequencing) as well as multiple taxonomic levels, data types, feature selection, and machine learning algorithms and evaluated the stressor prediction performance of each generated model for a total of 1,536 evaluated combinations of taxonomic datasets and data-processing methods. ResultsSequencing and data-processing methods had a substantial impact on stressor prediction. While omics-based methods detected much more taxa than amplicon sequencing, 16S sequencing outperformed all other sequencing methods in terms of stressor prediction based on the Matthews Correlation Coefficient. However, even the highest observed performance for 16S sequencing was still only moderate. Omics-based methods performed poorly overall, but this was likely due to insufficient sequencing depth. Data types had no impact on performance while feature selection significantly improved performance for omics-based methods but not for amplicon sequencing. ConclusionAmplicon sequencing might be a better candidate for machine-learning-based environmental stressor prediction than omics-based methods, but the latter require further research at higher sequencing depths to confirm this conclusion. More sampling could improve stressor prediction performance, and while this was not possible in the context of our study, thousands of sampling sites are monitored for routine environmental assessments, providing an ideal framework to further refine the approach for possible implementation in environmental diagnostics.

genomics↗