Search bioRxivSearch

Biology subjects

Balamotis, M.

Publications and source records attributed to Balamotis, M..

2 recordsLinked to original sources

Accurate Microbiome Sequencing with Synthetic Long Read Sequencing

The microbiome plays a central role in biochemical cycling and nutrient turnover of most ecosystems. Because it can comprise myriad microbial prokaryotes, eukaryotes and viruses, microbiome characterization requires high-throughput sequencing to attain an accurate identification and quantification of such co-existing microbial populations. Short-read next-generation-sequencing (srNGS) revolutionized the study of microbiomes and remains the most widely used approach, yet read lengths spanning only a few of the nine hypervariable regions of the 16S rRNA gene limit phylogenetic resolution leading to misclassification or failure to classify in a high percentage of cases. Here we evaluate a synthetic long-read (SLR) NGS approach for full-length 16S rRNA gene sequencing that is high-throughput, highly accurate and low-cost. The sequencing approach is amenable to highly multiplexed sequencing and provides microbiome sequence data that surpasses existing short and long-read modalities in terms of accuracy and phylogenetic resolution. We validated this commercially-available technology, termed LoopSeq, by characterizing the microbial composition of well-established mock microbiome communities and diverse real-world samples. SLR sequencing revealed differences in aquatic community complexity associated with environmental gradients, resolved species-level community composition of uterine lavage from subjects with histories of misconception and accurately detected strain differences, multiple copies of the 16S rRNA in a single strains genome, as well as low-level contamination in soil cyanobacterial cultures. This approach has implications for widespread adoption of high-resolution, accurate long-read microbiome sequencing as it is generated on popular short read sequencing platforms without the need for additional infrastructure.

microbiology

Targeted Transcriptome Analysis using Synthetic Long Read Sequencing Uncovers Isoform Reprograming in the Progression of Colon Cancer

Diversity in human gene expression stems, to a large extent, from splicing exons into multiple mRNA isoforms. Characterization of isoforms requires accurate long-read sequencing. However, read lengths, high error rates, low throughput and large input requirements are some of the challenges that remain to be addressed in sequencing technologies. In this study, we used a barcoding-based synthetic long read (SLR) isoform sequencing approach, LoopSeq, to generate sequencing reads sufficiently long and accurate to identify isoforms using standard short read Illumina sequencers. The method identifies isoforms from control RNA samples with 99.4% accuracy and a 0.01% per-base error rate, exceeding the accuracy reported for other long-read sequencing technologies. Applied to targeted transcriptome sequencing of over 10,000 genes from colon cancers and their metastatic counterparts, LoopSeq revealed large scale isoform redistributions from benign colon mucosa to primary colon cancer and metastatic cancer and identified several novel gene fusion isoforms in the colon cancer samples. Strikingly, our data showed that most single nucleotide variants (SNVs) occurred dominantly in specific isoforms and that some SNVs underwent isoform switching in cancer progression. The ability to use short read sequencers to generate accurate long-read isoform information as the raw unit of transcriptional information holds promise as a new and widely accessible approach in RNA isoform analyses.

genomics