Search bioRxivSearch

Biology subjects

Mahurkar, A.

Publications and source records attributed to Mahurkar, A..

4 recordsLinked to original sources

Cost Effective, Experimentally Robust Differential Expression Analysis for Human/Mammalian, Pathogen, and Dual-Species Transcriptomics

As sequencing read length has increased, researchers have quickly adopted longer reads for their experiments. Here, we examine host-pathogen interaction studies to assess if using longer reads is warranted. Six diverse datasets encountered in studies of host-pathogen interactions were used to assess what genomic attributes might affect the outcome of differential gene expression analysis including: gene density, operons, gene length, number of introns/exons, and intron length. Principal components analysis, hierarchical clustering with bootstrap support, and regression analyses of pairwise comparisons were undertaken on the same reads, looking at all combinations of paired and unpaired reads trimmed to 36,54,72, and 101-bp. For E coli, 36-bp single end reads performed as well as any other read length and as well as paired end reads. For all other comparisons, 54-bp and 72-bp reads were typically equivalent and different from 36-bp and 101-bp reads. Read pairing improved the outcome in several, but not all, comparisons in no discernable pattern, such that using paired reads is recommended in most scenarios. No specific genome attribute appeared to influence the data. However, experiments with an a priori expected greater biological complexity had more variable results with all read lengths relative to those with decreased complexity. When combined with cost, 54-bp paired end reads provided the most robust, internally reproducible results across all comparisons. However, using 36-bp single end reads may be desirable for bacterial samples, although possibly only if the transcriptional response is expected a priori to be robust.\n\nDATA SUMMARYO_LIThe human only CSHL Encode data set (1) was downloaded from ftp://hgdownload.cse.ucsc.edu/goldenPath/hgl9/encodeDCC/wgEncodeCshlLongRnaSeq/.\nC_LIO_LIThe data from mice vaginas infected with Candida albicans (2) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP057050).\nC_LIO_LIThe data from Aspergillus fumigatus cells in contact with human cells was downloaded from the SRA (url - https://www.ncbi.nlm.nih.gov/bioproject/399754).\nC_LIO_LIThe data from a strand-specific library from a study comparing C. albicans cells in contact with human cells with those in media (3) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP011085).\nC_LIO_LIThe data from C. albicans in culture media (3) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP011085).\nC_LIO_LIThe data from Escherichia coli grown in different media (4) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP056578).\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTAs sequencing technologies improve, sequencing costs decrease and read lengths increase. We examine host-pathogen interaction studies to assess if using these longer reads is warranted given their increased cost relative to using the same number of shorter reads. To this end we compared the use of various read lengths and read pairing for six diverse host-pathogen datasets with varying genomic attributes including: gene density, operons, gene length, number of introns/exons, and intron length. We find that in the bacterial sample, 36-bp single end reads performed as well as any other read length and as well as paired end reads. When combined with cost, 54-bp paired end reads provided the most robust, internally reproducible results for all other comparisons. Read pairing improved the outcome in several, but not all, comparisons in no discernable pattern, such that using paired reads is recommended in most scenarios. No specific genome attribute appeared to influence the data.

genomics

FADU: A Feature Counting Tool for Prokaryotic RNA-Seq Analysis

MotivationThe major algorithms for quantifying transcriptomics data for differential gene expression analysis were designed for analyzing data from human or human-like genomes, specifically those with single gene transcripts and distinct transcriptional boundaries that extend beyond the coding sequence (CDS) as identified through expressed sequence tags (ESTs) or EST-like sequence data. Some eukaryotic genomes and all, or nearly all, bacterial genomes require alternate methods of quantification since they lack annotation of transcriptional boundaries with EST or EST-like data, have overlapping transcriptional boundaries, and/or have polycistronic transcripts.\n\nResultsAn algorithm was developed and tested that better quantifies transcriptomics data for differential gene expression analysis in organisms with overlapping transcriptional units and polycistronic transcripts. Using data from standard libraries originating from Escherichia coli and Ehrlichia chaffeensis, and strand-specific libraries from the Wolbachia endosymbiont wBm, FADU can derive counts for genes that are missed by HTSeq and featurecounts. Using the default parameters with the E. coli data, FADU can detect transcription of 51 more genes than HTSeq in union mode and 21 genes more than featurecounts, with 42 and 18 of these features being <300 bp, respectively. Due to its ability to derive counts for otherwise unrepresented genes without overstating their abundance, we believe FADU to be an improved tool for quantifying transcripts in prokaryotic systems for RNA-Seq analyses.\n\nAvailability and implementationFADU is available at https://github.com/adkinsrs/FADU. FADU was implemented using Python3 and requires the PySAM module (version 0.12.0.1 or later).\n\nContactjdhotopp@som.umaryland.edu

genomics

Metaviz: interactive statistical and visual analysis of metagenomic data

Along with the survey techniques of 16S rRNA amplicon and whole-metagenome shotgun sequencing, an array of tools exists for clustering, taxonomic annotation, normalization, and statistical analysis of microbiome sequencing results. Integrative and interactive visualization that enables researchers to perform exploratory analysis in this feature rich hierarchical data is an area of need. In this work, we present Metaviz, a web browser-based tool for interactive exploratory metagenomic data analysis. Metaviz can visualize abundance data served from an R session or a Python web service that queries a graph database. As metagenomic sequencing features have a hierarchy, we designed a novel navigation mechanism to explore this feature space. We visualize abundance counts with heatmaps and stacked bar plots that are dynamically updated as a user selects taxonomic features to inspect. Metaviz also supports common data exploration techniques, including PCA scatter plots to interpret variability in the dataset and alpha diversity boxplots for examining ecological community composition. The Metaviz application and documentation is hosted at http://www.metaviz.org.

bioinformatics

Dual RNA sequencing (dRNA-Seq) of bacteria and their host cells

Bacterial pathogens subvert host cells by manipulating cellular pathways for survival and replication; in turn, host cells respond to the invading pathogen through cascading changes in gene expression. Deciphering these complex temporal and spatial dynamics to identify novel bacterial virulence factors or host response pathways is crucial for improved diagnostics and therapeutics. Dual RNA sequencing (dRNA-Seq) has recently been developed to simultaneously capture host and bacterial transcriptomes from an infected cell. This approach builds on the high sensitivity and resolution of RNA-Seq technology and is applicable to any bacteria that interact with eukaryotic cells, encompassing parasitic, commensal or mutualistic lifestyles. We pioneered dRNA-Seq to simultaneously capture prokaryotic and eukaryotic expression profiles of cells infected with bacteria, using in vitro Chlamydia-infected epithelial cells as proof of principle. Here we provide a detailed laboratory and bioinformatics protocol for dRNA-seq that is readily adaptable to any host-bacteria system of interest.

genomics