Search bioRxiv⌕ Search

Biology subjects

Schärfen, L.

Publications and source records attributed to Schärfen, L..

3 recordsLinked to original sources

Long-read Sequencing of Nascent RNA from Budding and Fission Yeasts

Gene expression requires DNA transcription and simultaneous RNA processing steps that transform the precursor RNA into fully mature RNA. In eukaryotes, the processing of protein-encoding messenger RNAs (mRNAs) includes 5 end capping, editing, splicing, RNA modification, poly-adenylation cleavage, and polyadenylation. Short-read sequencing of total or messenger RNA largely reveals the final output of transcription and processing because it utilizes 1) steady-state, mature RNA that is mostly processed and 2) sequencing reads that are too short to detect adjacent processing events (e.g. two adjacent introns). In contrast, long-read sequencing of nascent RNA allows the detection of rarer, full-length transcripts that are in the process of being transcribed and processed. The 3 end of each nascent RNA establishes the position of RNA polymerase II (Pol II) along the gene at the time of cell lysis, providing a timeline for RNA processing events. In addition, the density of 3 ends along genes or at gene landmarks reflects Pol II density, which is related to changes in transcription elongation rate. In organisms with complex gene architectures, information about splicing across multiple introns within the same transcript can be extracted, as well as the location of transcription start sites (TSSs) and polyA cleavage sites. Here, we describe the isolation of nascent RNA from the yeasts Saccharomyces cerevisiae and Schizosaccharomyces pombe, preparation of a cDNA library for long-read sequencing on Oxford Nanopore Technologies or Pacific Biosciences platforms, and initial data analysis steps. These methods comprise versatile and powerful tools for the investigation of coupled RNA synthesis and processing.

molecular biology↗

Rapid folding of nascent RNA regulates eukaryotic RNA biogenesis

An RNAs catalytic, regulatory, or coding potential depends on RNA structure formation. Because base pairing occurs during transcription, early structural states can govern RNA processing events and dictate the formation of functional conformations. These co-transcriptional states remain unknown. Here, we develop CoSTseq, which detects nascent RNA base pairing within and upon exit from RNA polymerases (Pols) transcriptome-wide in living yeast cells. By monitoring each nucleotides base pairing activity during transcription, we identify distinct classes of behaviors. While 47% of rRNA nucleotides remain unpaired, rapid and delayed base pairing - with rates of 48.5 and 13.2 kb-1 of transcribed rDNA, respectively - typically completes when Pol I is only 25 bp downstream. We show that helicases act immediately to remodel structures across the rDNA locus and facilitate ribosome biogenesis. In contrast, nascent pre-mRNAs attain local structures indistinguishable from mature mRNAs, suggesting that refolding behind elongating ribosomes resembles co-transcriptional folding behind Pol II.

biochemistry↗

Extensible benchmarking of methods that identify and quantify polyadenylation sites from RNA-seq data

The tremendous rate with which data is generated and analysis methods emerge makes it increasingly difficult to keep track of their domain of applicability, assumptions, and limitations and consequently, of the efficacy and precision with which they solve specific tasks. Therefore, there is an increasing need for benchmarks, and for the provision of infrastructure for continuous method evaluation. APAeval is an international community effort, organized by the RNA Society in 2021, to benchmark tools for the identification and quantification of the usage of alternative polyadenylation (APA) sites from short-read, bulk RNA-sequencing (RNA-seq) data. Here, we reviewed 17 tools and benchmarked eight on their ability to perform APA identification and quantification, using a comprehensive set of RNA-seq experiments comprising real, synthetic, and matched 3'-end sequencing data. To support continuous benchmarking, we have incorporated the results into the OpenEBench online platform, which allows for seamless extension of the set of methods, metrics, and challenges. We envisage that our analyses will assist researchers in selecting the appropriate tools for their studies. Furthermore, the containers and reproducible workflows generated in the course of this project can be seamlessly deployed and extended in the future to evaluate new methods or datasets.

bioinformatics↗