Search bioRxivSearch

Biology subjects

Spitale, R. C.

Publications and source records attributed to Spitale, R. C..

2 recordsLinked to original sources

Functional conservation of lncRNA JPX despite sequence and structural divergence

Long noncoding RNAs (lncRNAs) have been identified in all eukaryotes and are most abundant in the human genome. However, the functional importance and mechanisms of action for human lncRNAs are largely unknown. Using comparative sequence, structural, and functional analyses, we characterize the evolution and molecular function of human lncRNA JPX. We find that human JPX and its mouse homolog, lncRNA Jpx, have deep divergence in their nucleotide sequences and RNA secondary structures. Despite such differences, both lncRNAs demonstrate robust binding to CTCF, a protein that is central to Jpxs role in X chromosome inactivation. In addition, our functional rescue experiment using Jpx-deletion mutant cells, shows that human JPX can functionally complement the loss of Jpx in mouse embryonic stem cells. Our findings support a model for functional conservation of lncRNAs independent from sequence and structural changes. The study provides mechanistic insight into the evolution of lncRNA function.

molecular biology

A technology-agnostic long-read analysis pipeline for transcriptome discovery and quantification

Alternative splicing is widely acknowledged to be a crucial regulator of gene expression and is a key contributor to both normal developmental processes and disease states. While cost-effective and accurate for quantification, short-read RNA-seq lacks the ability to resolve full-length transcript isoforms despite increasingly sophisticated computational methods. Long-read sequencing platforms such as Pacific Biosciences (PacBio) and Oxford Nanopore (ONT) bypass the transcript reconstruction challenges of short reads. Here we introduce TALON, the ENCODE4 pipeline for platform-independent analysis of long-read transcriptomes. We apply TALON to the GM12878 cell line and show that while both PacBio and ONT technologies perform well at full-transcript discovery and quantification, each displayed distinct technical artifacts. We further apply TALON to mouse hippocampus and cortex transcriptomes and find that 422 genes found in these regions have more reads associated with novel isoforms than with annotated ones. We demonstrate that TALON is a capable of tracking both known and novel transcript models as well as their expression levels across datasets for both simple studies and in larger projects. These properties will enable TALON users to move beyond the limitations of short-read data to perform isoform discovery and quantification in a uniform manner on existing and future long-read platforms.

genomics