Search bioRxivSearch

Biology subjects

Hannah Dueck

Publications and source records attributed to Hannah Dueck.

2 recordsLinked to original sources

KimLabIDV: Application for Interactive RNA-Seq Data Analysis and Visualization

Many R packages have been developed for transcriptome analysis but their use often requires familiarity with R and integrating results of different packages is difficult. Here we present PIVOT, an R-based application with a uniform user interface and graphical data management that allows non-programmers to conveniently access various bioinformatics tools and interactively explore transcriptomics data. PIVOT supports many popular open source packages for transcriptome analysis and provides an extensive set of tools for statistical data manipulations. A graph-based visual interface is used to represent the links between derived datasets, allowing easy tracking of data versions. PIVOT further supports automatic report generation, publication-quality plots, and program/data state saving, such that all analysis can be saved, shared and reproduced.

Bioinformatics

IVT-seq reveals extreme bias in RNA-sequencing

BackgroundRNA sequencing (RNA-seq) is a powerful technique for identifying and quantifying transcription and splicing events, both known and novel. However, given its recent development and the proliferation of library construction methods, understanding the bias it introduces is incomplete but critical to realizing its value.\n\nResultsHere we present a method, in vitro transcription sequencing (IVT-seq), for identifying and assessing the technical biases in RNA-seq library generation and sequencing at scale. We created a pool of > 1000 in vitro transcribed (IVT) RNAs from a full-length human cDNA library and sequenced them with poly-A and total RNA-seq, the most common protocols. Because each cDNA is full length and we show IVT is incredibly processive, each base in each transcript should be equivalently represented. However, with common RNA-seq applications and platforms, we find [~]50% of transcripts have > 2-fold and [~]10% have > 10-fold differences in within-transcript sequence coverage. Strikingly, we also find > 6% of transcripts have regions of high, unpredictable sequencing coverage, where the same transcript varies dramatically in coverage between samples, confounding accurate determination of their expression. To get at causal factors, we used a combination of experimental and computational approaches to show that rRNA depletion is responsible for the most significant variability in coverage and that several sequence determinants also strongly influence representation.\n\nConclusionsIn sum, these results show the utility of IVT-seq in promoting better understanding of bias introduced by RNA-seq and suggest caution in its interpretation. Furthermore, we find that rRNA-depletion is responsible for substantial, unappreciated biases in coverage. Perhaps most importantly, these coverage biases introduced during library preparation suggest exon level expression analysis may be inadvisable.

Genomics