Search bioRxivSearch

Biology subjects

Andrew E Jaffe

Publications and source records attributed to Andrew E Jaffe.

10 recordsLinked to original sources

A framework for RNA quality correction in differential expression analysis

RNA sequencing (RNA-seq) is a powerful approach for measuring gene expression levels in cells and tissues, but it relies on high-quality RNA. We demonstrate here that statistical adjustment employing existing quality measures largely fails to remove the effects of RNA degradation when RNA quality associates with the outcome of interest. Using RNA-seq data from a molecular degradation experiment of human brain tissue, we introduce the quality surrogate variable (qSVA) analysis framework for estimating and removing the confounding effect of RNA quality in differential expression analysis. We show this approach results in greatly improved replication rates (>3x) across two large independent postmortem human brain studies of schizophrenia. Finally, we explored public datasets to demonstrate potential RNA quality confounding when comparing expression levels of different brain regions and diagnostic groups beyond schizophrenia. Our approach can therefore improve the interpretation of differential expression analysis of transcriptomic data from the human brain.

Bioinformatics

recount: A large-scale resource of analysis-ready RNA-seq expression data

recount is a resource of processed and summarized expression data spanning nearly 60,000 human RNA-seq samples from the Sequence Read Archive (SRA). The associated recount Bio-conductor package provides a convenient API for querying, downloading, and analyzing the data. Each processed study consists of meta/phenotype data, the expression levels of genes and their underlying exons and splice junctions, and corresponding genomic annotation. We also provide data summarization types for quantifying novel transcribed sequence including base-resolution coverage and potentially unannotated splice junctions. We present workflows illustrating how to use recount to perform differential expression analysis including meta-analysis, annotation-free base-level analysis, and replication of smaller studies using data from larger studies. recount provides a valuable and user-friendly resource of processed RNA-seq datasets to draw additional biological insights from existing public data. The resource is available at https://jhubiostatistics.shinyapps.io/recount/.

Genomics

Human splicing diversity across the Sequence Read Archive

We aligned 21,504 publicly available Illumina-sequenced human RNA-seq samples from the Sequence Read Archive (SRA) to the human genome and compared detected exon-exon junctions with junctions in several recent gene annotations. 56,865 junctions (18.6%) found in at least 1,000 samples were not annotated, and their expression associated with tissue type. Newer samples contributed few novel well-supported junctions, with 96.1% of junctions detected in at least 20 reads across samples present in samples before 2013. Junction data is compiled into a resource called intropolis available at http://intropolis.rail.bio. We discuss an application of this resource to cancer involving a recently validated isoform of the ALK gene.

Genomics

Strong Components of Epigenetic Memory in Cultured Human Fibroblasts Related to Site of Origin and Donor Age

Differentiating pluripotent cells from fibroblast progenitors is a potentially transformative tool in personalized medicine. We previously identified relatively greater success culturing dura-derived fibroblasts than scalp-derived fibroblasts from postmortem tissue. We hypothesized that these differences in culture success were related to epigenetic differences between the cultured fibroblasts by sampling location, and therefore generated genome-wide DNA methylation and transcriptome data on 11 intrinsically matched pairs of dural and scalp fibroblasts from donors across the lifespan (infant to 85 years). While these cultured fibroblasts were several generations removed from the primary tissue and morphologically indistinguishable, we found widespread epigenetic differences by sampling location at the single CpG (N=101,989), region (N=697), \"block\" (N=243), and global spatial scales suggesting a strong epigenetic memory of original fibroblast location. Furthermore, many of these epigenetic differences manifested in the transcriptome, particularly at the region-level. We further identified 7,265 CpGs and 11 regions showing significant epigenetic memory related to the age of the donor, as well as an overall increased epigenetic variability, preferentially in scalp-derived fibroblasts -83% of loci were more variable in scalp, hypothesized to result from cumulative exposure to environmental stimuli in the primary tissue. By integrating publicly available DNA methylation datasets on individual cell populations in blood and brain, we identified significantly increased inter-individual variability in our scalp- and other skin-derived fibroblasts on a similar scale as epigenetic differences between different lineages of blood cells. Lastly, these epigenetic differences did not appear to be driven by somatic mutation - while we identified 64 probable de-novo variants across the 11 subjects, there was no association between mutation burden and age of the donor (p=0.71). These results depict a strong component of epigenetic memory in cell culture from primary tissue, even after several generations of daughter cells, related to cell state and donor age.

Cell Biology

Rail-RNA: Scalable analysis of RNA-seq splicing and coverage

RNA sequencing (RNA-seq) experiments now span hundreds to thousands of samples. Current spliced alignment software is designed to analyze each sample separately. Consequently, no information is gained from analyzing multiple samples together, and it is difficult to reproduce the exact analysis without access to original computing resources. We describe Rail-RNA, a cloud-enabled spliced aligner that analyzes many samples at once. Rail-RNA eliminates redundant work across samples, making it more efficient as samples are added. For many samples, Rail-RNA is more accurate than annotation-assisted aligners. We use Rail-RNA to align 667 RNA-seq samples from the GEUVADIS project on Amazon Web Services in under 16 hours for US$0.91 per sample. Rail-RNA produces alignments and base-resolution bigWig coverage files, ready for use with downstream packages for reproducible statistical analysis. We identify expressed regions in the GEUVADIS samples and show that both annotated and unannotated (novel) expressed regions exhibit consistent patterns of variation across populations and with respect to known confounders. Rail-RNA is open-source software available at http://rail.bio.

Bioinformatics

regionReport: Interactive reports for region-based analyses

regionReport is a R package for generating detailed interactive reports from regions of the genome. The report includes quality-control checks, an overview of the results, an interactive table of the genomic regions, and reproducibility information. regionReport can easily be expanded with report templates for other specialized analyses. In particular, regionReport has an extensive report template for exploring derfinder results from annotation-agnostic RNA-seq differential expression analyses.\n\nAvailabilityregionReport is freely available via Bioconductor at bioconductor.org/packages/release/bioc/html/regionReport.html.

Bioinformatics

Flexible expressed region analysis for RNA-seq with derfinder

BackgroundDifferential expression analysis of RNA sequencing (RNA-seq) data typically relies on reconstructing transcripts or counting reads that overlap known gene structures. We previously introduced an intermediate statistical approach called differentially expressed region (DER) finder that seeks to identify contiguous regions of the genome showing differential expression signal at single base resolution without relying on existing annotation or potentially inaccurate transcript assembly.\n\nResultsWe present the derfinder software that improves our annotation-agnostic approach to RNA-seq analysis by: (1) implementing a computationally efficient bump-hunting approach to identify DERs which permits genome-scale analyses in a large number of samples, (2) introducing a flexible statistical modeling framework, including multi-group and time-course analyses and (3) introducing a new set of data visualizations for expressed region analysis. We apply this approach to public RNA-seq data from the Genotype-Tissue Expression (GTEx) project and BrainSpan project to show that derfinder permits the analysis of hundreds of samples at base resolution in R, identifies expression outside of known gene boundaries and can be used to visualize expressed regions at base-resolution. In simulations our base resolution approaches enable discovery in the presence of incomplete annotation and is nearly as powerful as feature-level methods when the annotation is complete.\n\nConclusionsderfinder analysis using expressed region-level and single base-level approaches provides a compromise between full transcript reconstruction and feature-level analysis.\n\nThe package is available from Bioconductor at www.bioconductor.org/packages/derfinder.

Bioinformatics

Polyester: simulating RNA-seq datasets with differential transcript expression

MotivationStatistical methods development for differential expression analysis of RNA sequencing (RNA-seq) requires software tools to assess accuracy and error rate control. Since true differential expression status is often unknown in experimental datasets, artificially-constructed datasets must be utilized, either by generating costly spike-in experiments or by simulating RNA-seq data.\n\nResultsPolyester is an R package designed to simulate RNA-seq data, beginning with an experimental design and ending with collections of RNA-seq reads. Its main advantage is the ability to simulate reads indicating isoform-level differential expression across biological replicates for a variety of experimental designs. Data generated by Polyester is a reasonable approximation to real RNA-seq data and standard differential expression workflows can recover differential expression set in the simulation by the user.\n\nAvailability and ImplementationPolyester is freely available from Bioconductor (http://bioconductor.org/).\n\nContactjtleek@gmail.com\n\nSupplementary InformationSupplementary figures are available online.

Bioinformatics

The methylome of the human frontal cortex across development

DNA methylation (DNAm) plays an important role in epigenetic regulation of gene expression, orchestrating tissue differentiation and development during all stages of mammalian life. This epigenetic control is especially important in the human brain, with extremely dynamic gene expression during fetal and infant life, and becomes progressively more stable at later periods of development. We characterized the epigenetic state of the developing and aging human frontal cortex in post-mortem tissue from 351 individuals across the lifespan using the Illumina 450k DNA methylation microarray. The largest changes in the methylome occur at birth at varying spatial resolutions - we identify 359,087 differentially methylated loci, which form 23,732 significant differentially methylated regions (DMRs). There were also 298 regions of long-range changes in DNAm, termed \"blocks\", associated with birth that strongly overlap previously published colon cancer \"blocks\". We then identify 55,439 DMRs associated with development and aging, of which 61.9% significantly associate with nearby gene expression levels. Lastly, we find enrichment of genomic loci of risk for schizophrenia and several other common diseases among these developmental DMRs. These data, integrated with existing genetic and transcriptomic data, create a rich genomic resource across brain development.

Developmental Biology

Flexible analysis of transcriptome assemblies with Ballgown

Introduction Introduction Negative control experiment Positive control experiment Confirmation of statistical... Analysis of RNA-seq experiments... Analysis of quantitative... Expression quantitative trait... Computational Efficiency Summary References A key advantage of RNA sequencing (RNA-seq) over hybridization-based technologies such as microarrays is that RNA-seq makes it possible to reconstruct complete gene structures, including multiple splice variants, from raw RNA-seq reads without relying on previously-established annotations [20, 32, 9]. But with this added flexibility, there are increased computational demands on upstream processing tasks such as alignment and ass ...

Bioinformatics