bioRxiv · 10.1101/195826
SAMSA2: A standalone metatranscriptome analysis pipeline
Abstract
BackgroundComplex microbial communities are an area of rapid growth in biology. Metatranscriptomics allows one to investigate the gene activity in an environmental sample via high-throughput sequencing. Metatranscriptomic experiments are computationally intensive because the experiments generate a large volume of sequence data and the sequences must be compared with many references.\n\nResultsHere we present SAMSA2, an upgrade to the original Simple Annotation of Metatranscriptomes by Sequence Analysis (SAMSA) pipeline that has been redesigned for use on a supercomputing cluster. SAMSA2 is faster due to the use of the DIAMOND aligner, and more flexible and reproducible because it uses local databases. SAMSA2 is available with detailed documentation, and example input and output files along with examples of master scripts for full pipeline execution.\n\nConclusionsUsing publicly available example data, we demonstrate that SAMSA2 is a rapid and efficient metatranscriptome pipeline for analyzing large paired-end RNA-seq datasets in a supercomputing cluster environment. SAMSA2 provides simplified output that can be examined directly or used for further analyses, and its reference databases may be upgraded, altered or customized to fit the specifics of any experiment.
Source connections
Explore related subjects
Keep this discovery
Westreich, S. T., Treiber, M. L., Mills, D. A., Korf, I., Lemay, D. G.. 2017-09-29. SAMSA2: A standalone metatranscriptome analysis pipeline. https://doi.org/10.1101/195826
Cite the original work for its findings. Save a collection to share your selection of sources.