Search bioRxiv⌕ Search

Biology subjects

Sim, A. D.

Publications and source records attributed to Sim, A. D..

4 recordsLinked to original sources

Isoform-level discovery, quantification and fusion analysis from single-cell and spatial long-read RNA-seq data with Bambu-Clump

Single cell and spatial transcriptomics have dramatically changed how we can profile RNA from heterogenous biological samples. Combining single cell and spatial profiling with long read RNA-Seq promises to enable the discovery and quantification of individual RNA isoforms at the single-cell level. However, highly multiplexed data such as from a single cell experiment only generates a limited number of reads for each cell, constituting a major challenge for transcript discovery and quantification with existing approaches that usually have limited power for samples with low sequencing depth. Here we present Bambu-Clump, a computational method that performs transcript discovery and quantification from single cell and spatial long read RNA-Seq data using information from both each cell and the cell cluster. Using this approach, Bambu-Clump provides the most accurate transcript discovery compared to other existing methods, and improves transcript quantification compared to methods that rely on estimates derived from single cells. We apply Bambu-Clump to identify fusion transcripts in single-cells, compare 5 and 3 selection protocols, and identify novel isoform cell-type markers in spatial mouse brain data. Together, Bambu-Clump provides an easy-to-use, efficient, and accurate method for analysing individual isoform expression for single cells and cell clusters across multiple datasets and replicates from long read RNA-Seq.

bioinformatics↗

CFC-seq: identification of full-length capped RNAs unveil enhancer-derived transcription

Long-read sequencing has transformed transcriptome profiling, yet capturing full-length, non-polyadenylated transcripts like enhancer RNAs (eRNAs) remains challenging. Here, we introduce CFC-seq, combining cap-trapping and in vitro poly(A)-tailing to sequence poly(A) and non-poly(A) RNAs with precise transcription start site. Paired with our assembler, SALA, we identified 39,425 novel transcriptional units, including [~]24,000 eRNAs. Our data reveal a distinct genomic code governing eRNA fate dictated by core promoter architecture. CpG-island enhancers show high chromatin connectivity but yield short, exosome-sensitive RNAs. Conversely, TATA-box enhancers systematically co-opt LTR retrotransposons to inherit structural motifs that produce long, stable, and spliced RNAs. Mechanistically, the pioneer factor NF-Y activates these viral elements to license transcription, balanced by TEAD4 activity across a dual-gear regulatory axis. Finally, non-poly(A) eRNAs terminate via exosome-associated processing at structural-depleted cleavage zones. This comprehensive annotation links enhancer sequence architecture to RNA fate, providing a new transformative framework for decoding the functional human genome. HighlightsO_LIExpanded genomic architecture: CFC-seq unmasks a hidden layer of human transcriptome, identifying 39,425 novel transcriptional units with high-confidence TSS support, including [~]24,000 eRNAs. C_LIO_LITSS-first assembler: We introduce SALA, a specialized long-read assembler that prioritizes authentic 5 Cap-trapped ends to accurately reconstruct the TSS-resolved transcript models. C_LIO_LIGenomic code of eRNA fate: CGI enhancers drive short and exosome-sensitive transcripts associated with repressive H3K27me3 mark and high chromatin connectivity. TATA-box enhancers produce cell-type-specific, long, stable, and frequently spliced eRNAs. C_LIO_LIEvolutionary co-option of retrotransposons: A major fraction of TATA-box eRNAs originate from LTR retrotransposons, providing a direct mechanism for integration of viral elements into the human regulatory landscape. C_LIO_LIA dual-gear pioneering axis: The pioneer factor NF-Y activates unprimed LTR-TATA enhancers to license transcription independent of histone acetylation cascades, operating in parallel with TEAD4-mediated activation. C_LIO_LIStructural determinants of eRNA termination: Non-poly(A) eRNA TES features a secondary structure depletion zone that coordinates pol II termination and calibrates exosome-mediated turnover. C_LI

genomics↗

Systematic assessment of long-read RNA-seq methods for transcript identification and quantification

AbstractThe Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.

genomics↗

Context-Aware Transcript Quantification from Long Read RNA-Seq data with Bambu

Most approaches to transcript quantification rely on fixed reference annotations. However, the transcriptome is dynamic, and depending on the context, such static annotations contain inactive isoforms for some genes while they are incomplete for others. To address this, we have developed Bambu, a method that performs machine-learning based transcript discovery to enable quantification specific to the context of interest using long-read RNA-Seq data. To identify novel transcripts, Bambu employs a precision-focused threshold referred to as the novel discovery rate (NDR), which replaces arbitrary per-sample thresholds with a single interpretable parameter. Bambu retains the full-length and unique read counts, enabling accurate quantification in presence of inactive isoforms. Compared to existing methods for transcript discovery, Bambu achieves greater precision without sacrificing sensitivity. We show that context-aware annotations improve abundance estimates for both novel and known transcripts. We apply Bambu to human embryonic stem cells to quantify isoforms from repetitive HERVH-LTR7 retrotransposons, demonstrating the ability to estimate transcript expression specific to the context of interest.

bioinformatics↗