Search bioRxiv⌕ Search

Biology subjects

Nabizadehmashhadtoroghi, M.

Publications and source records attributed to Nabizadehmashhadtoroghi, M..

2 recordsLinked to original sources

Messenger-RNA Modification Standards and Machine Learning Models Facilitate Absolute Site-Specific Pseudouridine Quantification

Enzyme-mediated chemical modifications to mRNA are important for fine-tuning gene expression, but they are challenging to quantify due to low copy number and limited tools for accurate detection. Existing studies have typically focused on the identification and impact of adenine modifications on mRNA (m6A and inosine) due to the availability of analytical methods. The pseudouridine ({Psi}) mRNA modification is also highly abundant but difficult to detect and quantify because there is no available antibody, it is mass silent, and maintains canonical basepairing with adenine. Nanopores may be used to directly identify {Psi} sites in RNAs using a systematically miscalled base, however, this approach is not quantitative and highly sequence dependent. In this work, we apply supervised machine learning models that are trained on sequence-specific, synthetic controls to endogenous transcriptome data and achieve the first quantitative {Psi} occupancy measurement in human mRNAs. Our supervised machine learning models reveal that for every site studied, different signal parameters are required to maximize {Psi} classification accuracy. We show that applying our model is critical for quantification, especially in low-abundance mRNAs. Our engine can be used to profile {Psi}-occupancy across cell types and cell states, thus providing critical insights about physiological relevance of {Psi} modification to mRNAs.

genomics↗

Detection of pseudouridine modifications and type I/II hypermodifications in human mRNAs using direct, long-read sequencing

We developed and applied a semi-quantitative method for high-confidence identification of pseudouridylated sites on mammalian mRNAs via direct long-read nanopore sequencing. A comparative analysis of a modification-free transcriptome reveals that the depth of coverage and specific k-mer sequences are critical parameters for accurate basecalling. By adjusting these parameters for high-confidence U-to-C basecalling errors, we identified many known sites of pseudouridylation and uncovered new uridine-modified sites, many of which fall in k-mers that are known targets of pseudouridine synthases. Identified sites were validated using 1,000-mer synthetic RNA controls bearing a single pseudouridine in the center position which demonstrate systematical under-calling using our approach. We identify mRNAs with up to 7 unique modification sites. Our pipeline allows direct detection of low-, medium-, and high-occupancy pseudouridine modifications on native RNA molecules from nanopore sequencing data as well as multiple modifications on the same strand.

bioengineering↗