Search bioRxiv⌕ Search

Biology subjects

Nguyen, A. N. T.

Publications and source records attributed to Nguyen, A. N. T..

2 recordsLinked to original sources

Customizable host and viral transcript enrichment using CRISPR-Cas9 long-read sequencing for isoform discovery and validation

Long-read RNA sequencing has been broadly utilized to examine the diversity of transcriptomes, understand differential expression and discover novel transcript isoforms. One of the major limitations of whole transcriptome sequencing is the difficulty in obtaining sufficient depth for low abundant transcripts. Methods which address this are either difficult to scale or customize: long- range PCR is customizable but difficult to scale beyond a few targets; probe hybridization panels are suited for scaling but require substantial investment to customize. In this study, we adopted RNA-guided CRISPR-Cas9 nuclease-based enrichment to target specific human and SARS-CoV-2 transcripts followed by long-read sequencing, utilizing minimal number of guide RNAs per target isoform. Our findings demonstrate that the CRISPR-Cas system is a highly effective method for customizable long-read sequencing of target transcripts while maintaining the accuracy of relative gene expression levels. The results highlight a valuable method for future research on transcript enrichment for isoform identification and low abundance transcript detection in infectious disease diagnosis.

genomics↗

Benchmarking reveals superiority of deep learning variant callers on bacterial nanopore sequence data

Variant calling is fundamental in bacterial genomics, underpinning the identification of disease transmission clusters, the construction of phylogenetic trees, and antimicrobial resistance prediction. This study presents a comprehensive benchmarking of SNP and indel variant calling accuracy across 14 diverse bacterial species using Oxford Nanopore Technologies (ONT) and Illumina sequencing. We generate gold standard reference genomes and project variations from closely-related strains onto them, creating biologically realistic distributions of SNPs and indels. Our results demonstrate that ONT variant calls from deep learning-based tools delivered higher SNP and indel accuracy than traditional methods and Illumina, with Clair3 providing the most accurate results overall. We investigate the causes of missed and false calls, highlighting the limitations inherent in short reads and discover that ONTs traditional limitations with homopolymer-induced indel errors are absent with high-accuracy basecalling models and deep learning-based variant calls. Furthermore, our findings on the impact of read depth on variant calling offer valuable insights for sequencing projects with limited resources, showing that 10x depth is sufficient to achieve variant calls that match or exceed Illumina. In conclusion, our research highlights the superior accuracy of deep learning tools in SNP and indel detection with ONT sequencing, challenging the primacy of short-read sequencing. The reduction of systematic errors and the ability to attain high accuracy at lower read depths enhance the viability of ONT for widespread use in clinical and public health bacterial genomics.

bioinformatics↗