Search bioRxivSearch

Biology subjects

Corrada Bravo, H.

Publications and source records attributed to Corrada Bravo, H..

7 recordsLinked to original sources

Assessing 16S marker gene survey data analysis methods using mixtures of human stool sample DNA extracts.

BackgroundAnalysis of 16S rRNA marker-gene surveys, used to characterize prokaryotic microbial communities, may be performed by numerous bioinformatic pipelines and downstream analysis methods. However, there is limited guidance on how to decide between methods, appropriate data sets and statistics for assessing these methods are needed. We developed a mixture dataset with real data complexity and an expected value for assessing 16S rRNA bioinformatic pipelines and downstream analysis methods. We generate an assessment dataset using a two-sample titration mixture design. The sequencing data were processed using multiple bioinformatic pipelines, i) DADA2 a sequence inference method, ii) Mothur a de novo clustering method, and iii) QIIME with open-reference clustering. The mixture dataset was used to qualitatively and quantitatively assess count tables generated using the pipelines.\n\nResultsThe qualitative assessment was used to evalute features only present in unmixed samples and titrations. The abundance of Mothur and QIIME features specific to unmixed samples and titrations were explained by sampling alone. However, for DADA2 over a third of the unmixed sample and titration specific feature abundance could not be explained by sampling alone. The quantitative assessment evaluated pipeline performance by comparing observed to expected relative and differential abundance values. Overall the observed relative abundance and differential abundance values were consistent with the expected values. Though outlier features were observed across all pipelines.\n\nConclusionsUsing a novel mixture dataset and assessment methods we quantitatively and qualitatively evaluated count tables generated using three bioinformatic pipelines. The dataset and methods developed for this study will serve as a valuable community resource for assessing 16S rRNA marker-gene survey bioinformatic methods.

bioinformatics

Fast and interpretable alternative splicing and differential gene-level expression analysis using transcriptome segmentation with Yanagi

IntroductionAnalysis of differential alternative splicing from RNA-seq data is complicated by the fact that many RNA-seq reads map to multiple transcripts, besides, the annotated transcripts are often a small subset of the possible transcripts of a gene. Here we describe Yanagi, a tool for segmenting transcriptome to create a library of maximal L-disjoint segments from a complete transcriptome annotation. That segment library preserves all transcriptome substrings of length L and transcripts structural relationships while eliminating unnecessary sequence duplications.\n\nContributionsIn this paper, we formalize the concept of transcriptome segmentation and propose an efficient algorithm for generating segment libraries based on a length parameter dependent on specific RNA-Seq library construction. The resulting segment sequences can be used with pseudo-alignment tools to quantify expression at the segment level. We characterize the segment libraries for the reference transcriptomes of Drosophila melanogaster and Homo sapiens and provide gene-level visualization of the segments for better interpretability. Then we demonstrate the use of segments-level quantification into gene expression and alternative splicing analysis. The notion of transcript segmentation as introduced here and implemented in Yanagi opens the door for the application of lightweight, ultra-fast pseudo-alignment algorithms in a wide variety of RNA-seq analyses.\n\nConclusionUsing segment library rather than the standard transcriptome succeeds in significantly reducing ambigious alignments where reads are multimapped to several sequences in the reference. That allowed avoiding the quantification step required by standard kmer-based pipelines for gene expression analysis. Moreover, using segment counts as statistics for alternative splicing analysis enables achieving comparable performance to counting-based approaches (e.g. rMATS) while rather using fast and lighthweight pseudo alignment.

bioinformatics

metagenomeFeatures: An R package for working with 16S rRNA reference databases and marker-gene survey feature data.

We developed the metagenomeFeatures R Bioconductor package along with annotation packages for the three primary 16S rRNA databases (Greengenes, RDP, and SILVA) to facilitate working with 16S rRNA sequence databases and marker-gene survey feature data. The metagenomeFeatures package defines two classes, MgDb for working with 16S rRNA sequence databases, and mgFeatures for working with marker-gene survey feature data. The associated annotation packages provide a consistent interface to the different 16S rRNA databases facilitating database comparison and exploration. The mgFeatures represents a crucial step in the development of a common data structure for working with 16S marker-gene survey data in R.\n\nAvailabilityhttps://bioconductor.org/packages/release/bioc/html/metagenomeFeatures.html\n\nContactnolson@nist.gov

bioinformatics

Analysis And Correction Of Compositional Bias In Sparse Sequencing Count Data

Count data derived from high-throughput DNA sequencing is frequently used in quantitative molecular assays. Due to properties inherent to the sequencing process, unnormalized count data is compositional, measuring relative and not absolute abundances of the assayed features. This compositional bias confounds inference of absolute abundances. We demonstrate that existing techniques for estimating compositional bias fail with sparse metagenomic 16S count data and propose an empirical Bayes normalization approach to overcome this problem. In addition, we clarify the assumptions underlying frequently used scaling normalization methods in light of compositional bias, including scaling methods that were not designed directly to address it.

genomics

Metaviz: interactive statistical and visual analysis of metagenomic data

Along with the survey techniques of 16S rRNA amplicon and whole-metagenome shotgun sequencing, an array of tools exists for clustering, taxonomic annotation, normalization, and statistical analysis of microbiome sequencing results. Integrative and interactive visualization that enables researchers to perform exploratory analysis in this feature rich hierarchical data is an area of need. In this work, we present Metaviz, a web browser-based tool for interactive exploratory metagenomic data analysis. Metaviz can visualize abundance data served from an R session or a Python web service that queries a graph database. As metagenomic sequencing features have a hierarchy, we designed a novel navigation mechanism to explore this feature space. We visualize abundance counts with heatmaps and stacked bar plots that are dynamically updated as a user selects taxonomic features to inspect. Metaviz also supports common data exploration techniques, including PCA scatter plots to interpret variability in the dataset and alpha diversity boxplots for examining ecological community composition. The Metaviz application and documentation is hosted at http://www.metaviz.org.

bioinformatics

Longitudinal differential abundance analysis of microbial marker-gene surveys using smoothing splines

BackgroundHigh-throughput targeted sequencing of the 16S ribosomal RNA marker gene is often used to profile and characterize the taxonomic composition of microbial communities. This type of big high-through sequencing data is rapidly being applied to various infectious diseases like diarrhea. While many studies are limited to single \"snapshots\" of these communities, there is increasing recognition that longitudinal profiling of these communities are required to understand community dynamics and the complex relationships between dynamics and phenotypes of interest. Statistical methods that determine microbial features that are differentially expressed are required as an initial step to characterizing phenotypic associations with community dynamics in big data and infectious diseases.\n\nResultsWe present a novel method for longitudinal marker-gene surveys based on smoothing splines that allows discovery and inference of time periods where specific microbial features are differentially abundant. We applied our method to three 16S marker-gene surveys, including, groups of gnotobiotic mice on two diets, patients challenged with ETEC (H10407), and a vaginal microbiome of healthy women. Employing our methodology we recover known bacterial differences and highlight a few extra species providing insight into when specific changes occurred. Additionally, in the cohort challenged with ETEC we recover proposed probiotic bacteria Bacteroides xylanisolvens, Collinsella aerofaciens, and Faecalibacterium prausnitzii associatons with healthy individuals.\n\nConclusionsThe method presented is, to our knowledge, the first flexible method of its kind implemented as a software capable of detecting time periods of differential abundance for microbial features species between two or more sample groups of interest. Our method is available within the metagenomeSeq open-source software for analysis of metagenomic package available through the Bioconductor project and is termed metaSplines.

genomics

Smooth Quantile Normalization

Between-sample normalization is a critical step in genomic data analysis to remove systematic bias and unwanted technical variation in high-throughput data. Global normalization methods are based on the assumption that observed variability in global properties is due to technical reasons and are unrelated to the biology of interest. For example, some methods correct for differences in sequencing read counts by scaling features to have similar median values across samples, but these fail to reduce other forms of unwanted technical variation. Methods such as quantile normalization transform the statistical distributions across samples to be the same and assume global differences in the distribution are induced by only technical variation. However, it remains unclear how to proceed with normalization if these assumptions are violated, for example if there are global differences in the statistical distributions between biological conditions or groups, and external information, such as negative or control features, is not available. Here we introduce a generalization of quantile normalization, referred to as smooth quantile normalization (qsmooth), which is based on the assumption that the statistical distribution of each sample should be the same (or have the same distributional shape) within biological groups or conditions, but allowing that they may differ between groups. We illustrate the advantages of our method on several high-throughput datasets with global differences in distributions corresponding to different biological conditions. We also perform a Monte Carlo simulation study to illustrate the bias-variance tradeoff of qsmooth compared to other global normalization methods. A software implementation is available from https://github.com/stephaniehicks/qsmooth.

genomics