Search bioRxivSearch

Biology subjects

Kasper D Hansen

Publications and source records attributed to Kasper D Hansen.

5 recordsLinked to original sources

recount: A large-scale resource of analysis-ready RNA-seq expression data

recount is a resource of processed and summarized expression data spanning nearly 60,000 human RNA-seq samples from the Sequence Read Archive (SRA). The associated recount Bio-conductor package provides a convenient API for querying, downloading, and analyzing the data. Each processed study consists of meta/phenotype data, the expression levels of genes and their underlying exons and splice junctions, and corresponding genomic annotation. We also provide data summarization types for quantifying novel transcribed sequence including base-resolution coverage and potentially unannotated splice junctions. We present workflows illustrating how to use recount to perform differential expression analysis including meta-analysis, annotation-free base-level analysis, and replication of smaller studies using data from larger studies. recount provides a valuable and user-friendly resource of processed RNA-seq datasets to draw additional biological insights from existing public data. The resource is available at https://jhubiostatistics.shinyapps.io/recount/.

Genomics

Whole genome analysis of the methylome and hydroxymethylome in normal and malignant lung and liver

DNA methylation at the 5-postion of cytosine (5mC) is a well-established epigenetic modification which regulates gene expression and cellular plasticity in development and disease. The ten-eleven translocation (TET) gene family is able to oxidize 5mC to 5-hydroxymethyl-cytosine (5hmC), providing an active mechanism for DNA demethylation, and may also provide its own regulatory function. Here we applied oxidative bisulfite sequencing to generate whole-genome DNA methylation and hydroxymethylation maps at single-base resolution in paired human liver and lung normal and cancer. We found that 5hmC is significantly enriched in CpG island (CGI) shores while depleted in CGIs themselves, in particular at active genes, resulting in a 5hmC but not 5mC bimodal distribution around CGI corresponding to H3K4me1 marks. Hydroxymethylation on promoters, gene bodies, and transcription termination regions showed strong positive correlation with gene expression within and across tissues, suggesting that 5hmC is a mark of active genes and could play a role gene expression mediated by DNA demethylation. Comparative analysis of methylomes and hydroxymethylomes revealed that 5hmC is significantly enriched in both tissue specific DMRs (t-DMRs) and cancer specific DMRs (c-DMRs), and 5hmC is negatively correlated with methylation changes, particularly in non-CGI associated DMRs. Together these findings indicate that changes in 5mC as well as in 5hmC and coupled to H3K4me1 correspond to differential gene expression in tissues and matching tumors, revealing an intricate gene expression regulation through interplay of methylome, hydroxyl-methylome, and histone modifications.

Genomics

Human splicing diversity across the Sequence Read Archive

We aligned 21,504 publicly available Illumina-sequenced human RNA-seq samples from the Sequence Read Archive (SRA) to the human genome and compared detected exon-exon junctions with junctions in several recent gene annotations. 56,865 junctions (18.6%) found in at least 1,000 samples were not annotated, and their expression associated with tissue type. Newer samples contributed few novel well-supported junctions, with 96.1% of junctions detected in at least 20 reads across samples present in samples before 2013. Junction data is compiled into a resource called intropolis available at http://intropolis.rail.bio. We discuss an application of this resource to cancer involving a recently validated isoform of the ALK gene.

Genomics

Rail-dbGaP: analyzing dbGaP-protected data in the cloud with Amazon Elastic MapReduce

Motivation: Public archives contain thousands of trillions of bases of valuable sequencing data. More than 40% of the Sequence Read Archive is human data protected by provisions such as dbGaP To analyze dbGaP-protected data, researchers must typically work with IT administrators and signing officials to ensure all levels of security are implemented at their institution. This is a major obstacle, impeding reproducibility and reducing the utility of archived data.\n\nResults: We present a protocol and software tool for analyzing protected data in a commercial cloud. The protocol, Rail-dbGaP, is applicable to any tool running on Amazon Web Services Elastic MapReduce. The tool, Rail-RNA v0.2, is a spliced aligner for RNA- seq data, which we demonstrate by running on 9,662 samples from the dbGaP-protected GTEx consortium dataset. The Rail-dbGaP protocol makes explicit for the first time the steps an investigator must take to develop Elastic MapReduce pipelines that analyze dbGaP-protected data in a manner compliant with NIH guidelines. Rail-RNA automates implementation of the protocol, making it easy for typical biomedical investigators to study protected RNA-seq data, regardless of their local IT resources or expertise.\n\nAvailability: Rail-RNA is available from http://rail.bio. Technical details on the Rail-dbGaP protocol as well as an implementation walkthrough are available at https://github.com/nellore/rail-dbgap. Detailed instructions on running Rail-RNA on dbGaP-protected data using Amazon Web Services are available at http://docs.rail.bio/dbgap/.\n\nContact: anellore@gmail.com, langmea@cs.jhu.edu

Bioinformatics

Reconstructing A/B compartments as revealed by Hi-C using long-range correlations in epigenetic data

Analysis of Hi-C data has shown that the genome can be divided into two compartments called A/B compartments. These compartments are cell-type specific and are associated with open and closed chromatin. We show that A/B compartments can be reliably estimated using epigenetic data from several different platforms, the Illumina 450k DNA methylation microarray, DNase hypersensitivity sequencing, single-cell ATAC sequencing and single-cell whole-genome bisulfite sequencing. We do this by exploiting the fact that the structure of long range correlations differs between open and closed compartments. This work makes A/B compartments readily available in a wide variety of cell types, including many human cancers.

Genomics