Search bioRxivSearch

Biology subjects

Elizabeth Tseng

Publications and source records attributed to Elizabeth Tseng.

2 recordsLinked to original sources

HapIso : An Accurate Method for the Haplotype-Specific Isoforms Reconstruction from Long Single-Molecule Reads

Sequencing of RNA provides the possibility to study an individuals transcriptome landscape and determine allelic expression ratios. Single-molecule protocols generate multi-kilobase reads longer than most transcripts allowing sequencing of complete haplotype isoforms. This allows partitioning the reads into two parental haplotypes. While the read length of the single-molecule protocols is long, the relatively high error rate limits the ability to accurately detect the genetic variants and assemble them into the haplotype-specific isoforms. In this paper, we present HapIso (Haplotype-specific Isoform Reconstruction), a method able to tolerate the relatively high error-rate of the single-molecule platform and partition the isoform reads into the parental alleles. Phasing the reads according to the allele of origin allows our method to efficiently distinguish between the read errors and the true biological mutations. HapIso uses a k-means clustering algorithm aiming to group the reads into two meaningful clusters maximizing the similarity of the reads within cluster and minimizing the similarity of the reads from different clusters. Each cluster corresponds to a parental haplotype. We use family pedigree information to evaluate our approach. Experimental validation suggests that HapIso is able to tolerate the relatively high error-rate and accurately partition the reads into the parental alleles of the isoform transcripts. Furthermore, our method is the first method able to reconstruct the haplotype-specific isoforms from long single-molecule reads.\n\nThe open source Python implementation of HapIso is freely available for download at https://github.com/smangul1/HapIso/

Bioinformatics

Widespread polycistronic transcripts in mushroom-forming fungi revealed by single-molecule long-read mRNA sequencing

Genes in prokaryotic genomes are often arranged into clusters and co-transcribed into polycistronic RNAs. Isolated examples of polycistronic RNAs were also reported in some eukaryotes but their presence was generally considered rare. Here we developed a long-read sequencing strategy to identify polycistronic transcripts in several mushroom forming fungal species including Plicaturopsis crispa, Phanerochaete chrysosporium, Trametes versicolor and Gloeophyllum trabeum1. We found genome-wide prevalence of polycistronic transcription in these Agaricomycetes, and it involves up to 8% of the transcribed genes. Unlike polycistronic mRNAs in prokaryotes, these co-transcribed genes are also independently transcribed, and upstream transcription may interfere downstream transcription. Further comparative genomic analysis indicates that polycistronic transcription is likely a feature unique to these fungi. In addition, we also systematically demonstrated that short-read assembly is insufficient for mRNA isoform discovery, especially for isoform-rich loci. In summary, our study revealed, for the first time, the genome prevalence of polycistronic transcription in a subset of fungi. Futhermore, our long-read sequencing approach combined with bioinformatics pipeline is a generic powerful tool for precise characterization of complex transcriptomes.

Genomics