Search bioRxiv⌕ Search

Biology subjects

Esfahani, N. G.

Publications and source records attributed to Esfahani, N. G..

3 recordsLinked to original sources

tRNA isodecoder analysis using Nanopore ionic current signals and deep learning

tRNA are short non-coding RNA characterized by their distinct tertiary structure and abundant chemical modifications. Conventional analysis strategies do not fully characterize tRNA isodecoders. We demonstrate that this limitation can be resolved for tRNA using nanopore ionic current data. We developed tRNAZAP, a deep learning strategy that uses nanopore ionic current signal information to classify native tRNA strands at isodecoder-level resolution without relying on sequence information. Additionally, the ionic current level classification allows for pairwise alignment of read sequences to reference sequences, producing optimal tRNA alignments. We applied tRNAZAP to direct tRNA sequencing data from Escherichia coli and Saccharomyces cerevisiae, and recovered 2.6% and 13.1% more aligned reads than BWA-MEM, respectively. tRNAZAP resolved these reads at an isodecoder-level and with consistently higher alignment identity. tRNAZAP is a powerful complement to sequence-based profiling and can contribute towards resolving the isodecoder landscape in more complex organisms including humans.

genomics↗

Evaluation of Nanopore direct RNA sequencing updates for modification detection

Nanopore technology can directly sequence intact RNA molecules, offering a unique capability to read native modifications. Oxford Nanopore Technologies recently updated its direct RNA sequencing technology from RNA002 to RNA004 chemistry. This update included an improved basecaller (Dorado) for increased sequencing accuracy, and ionic current models for de novo detection of four RNA modifications. Using a single RNA extraction from GM12878 B-lymphocyte cell line, we compared RNA002 and RNA004 sequencing chemistries and evaluated Dorado modification calling accuracy. We computed U-to-C mismatches, previously used to identify putative pseudouridine sites, and ran m6anet for identifying putative N6-methyladenosine sites. Dorado results for each respective modification showed both global and site-specific differences when compared to RNA002 results. We used DRS data from in vitro transcription of GM12878 genomic DNA as well as synthetic oligonucleotides to evaluate Dorado modification calling performance. Dorados pseudouridine model achieved 96-98% for both accuracy and F1-score. Similarly, Dorados N6-methyladenosine model achieved 94-98% accuracy, 96-99% F1-score. Our results demonstrate that Nanopore Direct RNA sequencing could simultaneously detect pseudouridine, N6-methyladenosine, 5-methylcytosine, and inosine modifications on individual mRNA strands. It is critical to validate these calls using orthogonal methods (e.g., Liquid Chromatography with Tandem Mass Spectrometry) for increased confidence. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC="FIGDIR/small/651717v1_ufig1.gif" ALT="Figure 1"> View larger version (25K): org.highwire.dtl.DTLVardef@1215534org.highwire.dtl.DTLVardef@160f38forg.highwire.dtl.DTLVardef@164a9corg.highwire.dtl.DTLVardef@17c72e5_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗

Genomic in vitro transcription and Nanopore direct RNA sequencing of a human B-Lymphocyte cell line

Genomic DNA used as a template for in vitro transcription of RNA can serve as a true negative control for benchmarking RNA modification detection by Nanopore direct RNA sequencing (DRS) models. We generated DRS data for in vitro transcribed (IVT) RNA composed of canonical nucleotides using genomic DNA from a human cell line. We applied Dorado modification calling models to these data, and calculated 9-mer specific false-positive rates for eight RNA modifications as a comparison point for future development of RNA modification models. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=119 SRC="FIGDIR/small/650674v2_ufig1.gif" ALT="Figure 1"> View larger version (23K): org.highwire.dtl.DTLVardef@962feorg.highwire.dtl.DTLVardef@4237aborg.highwire.dtl.DTLVardef@154d346org.highwire.dtl.DTLVardef@1fad8f2_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗