Search bioRxivSearch

Biology subjects

Carl Kingsford

Publications and source records attributed to Carl Kingsford.

4 recordsLinked to original sources

A pathway-centric view of spatial proximity in the 3D nucleome across cell lines

Spatial organization of the genome is critical for condition-specific gene expression. Previous studies have shown that functionally related genes tend to be spatially proximal. However, these studies have not been extended to multiple human cell types, and the extent to which context-specific spatial proximity of a pathway is related to its context-specific activity is not known. We report the first pathway-centric analyses of spatial proximity in six human cell lines. We find that spatial proximity of genes in a pathway tends to be context-specific, in a manner consistent with the pathways context-specific expression and function; housekeeping genes are ubiquitously proximal to each other, and cancer-related pathways such as p53 signaling are uniquely proximal in hESC. Intriguingly, we find a correlation between the spatial proximity of genes and interactions of their protein products, even after accounting for the propensity of co-pathway proteins to interact. Related pathways are also often spatially proximal to one another, and housekeeping genes tend to be proximal to several other pathways suggesting their coordinating role. Further, the spatially proximal genes in a pathway tend to be the drivers of the pathway activity and are enriched for transcription, splicing and transport functions. Overall, our analyses reveal a pathway-centric organization of the 3D nucleome whereby functionally related and interacting genes, particularly the initial drivers of pathway activity, but also genes across multiple related pathways, are in spatial proximity in a context-specific way. Our results provide further insights into the role of differential spatial organization in cell type-specific pathway activity.

Bioinformatics

Salmon provides accurate, fast, and bias-aware transcript expression estimates using dual-phase inference

We introduce Salmon, a new method for quantifying transcript abundance from RNA-seq reads that is highly-accurate and very fast. Salmon is the first transcriptome-wide quantifier to model and correct for fragment GC content bias, which we demonstrate substantially improves the accuracy of abundance estimates and the reliability of subsequent differential expression analysis compared to existing methods that do not account for these biases. Salmon achieves its speed and accuracy by combining a new dual-phase parallel inference algorithm and feature-rich bias models with an ultra-fast read mapping procedure. These innovations yield both exceptional accuracy and order-of-magnitude speed benefits over alignment-based methods.

Bioinformatics

Isoform-level Ribosome Occupancy Estimation Guided by Transcript Abundance with Ribomap

Ribosome profiling is a recently developed high-throughput sequencing technique that captures approximately 30 bp long ribosome-protected mRNA fragments during translation. Because of alternative splicing and repetitive sequences, a ribosome-protected read may map to many places in the transcriptome, leading to discarded or arbitrary mappings when standard approaches are used. We present a technique and software that addresses this problem by assigning reads to potential origins proportional to estimated transcript abundance. This yields a more accurate estimate of ribosome profiles compared with a na{iota}ve mapping. Ribomap is available as open source at http://www.cs.cmu.edu/[~]ckingsf/software/ribomap.

Bioinformatics

Compression of short-read sequences using path encoding

Storing, transmitting, and archiving the amount of data produced by next generation sequencing is becoming a significant computational burden. For example, large-scale RNA-seq meta-analyses may now routinely process tens of terabytes of sequence. We present here an approach to biological sequence compression that reduces the difficulty associated with managing the data produced by large-scale transcriptome sequencing. Our approach offers a new direction by sitting between pure reference-based compression and reference-free compression and combines much of the benefit of reference-based approaches with the flexibility of de novo encoding. Our method, called path encoding, draws a connection between storing paths in de Bruijn graphs --- a common task in genome assembly --- and context-dependent arithmetic coding. Supporting this method is a system, called a bit tree, to compactly store sets of kmers that is of independent interest. Using these techniques, we are able to encode RNA-seq reads using 3% -- 11% of the space of the sequence in raw FASTA files, which is on average more than 34% smaller than recent competing approaches. We also show that even if the reference is very poorly matched to the reads that are being encoded, good compression can still be achieved.

Bioinformatics