Search bioRxivSearch

Biology subjects

Fufa, T. D.

Publications and source records attributed to Fufa, T. D..

2 recordsLinked to original sources

Building the Mega Single Cell Transcriptome Ocular Meta-Atlas

The development of highly scalable single cell transcriptome technology has resulted in the creation of thousands of datasets, over 30 in the retina alone. Analyzing the transcriptomes between different projects is highly desirable as this would allow for better assessment of which biological effects are consistent across independent studies. However it is difficult to compare and contrast data across different projects as there are substantial batch effects from computational processing, single cell technology utilized, and the natural biological variation. While many single cell transcriptome specific batch correction methods purport to remove the technical noise it is difficult to ascertain which method functions works best. We developed a lightweight R package (scPOP) that brings in batch integration methods and uses a simple heuristic to balance batch merging and celltype/cluster purity. We use this package along with a Snakefile based workflow system to demonstrate how to optimally merge 766,615 cells from 33 retina datsets and three species to create a massive ocular single cell transcriptome meta-atlas. This provides a model how to efficiently create meta-atlases for tissues and cells of interest.

bioinformatics

A short read de novo transcriptome construction pipeline optimized by long reads reveals novel developmentally regulated gene isoforms and disease targets in hundreds of eye samples

De novo transcriptome construction from short-read RNA-seq is a common method for reconstructing mRNA transcripts within a given sample. However, the precision of this process is unclear as it is difficult to obtain a ground-truth measure of transcript expression. With advances in third generation sequencing, full length transcripts of whole transcriptomes can be accurately sequenced to generate a ground-truth transcriptome. We generated long-read PacBio and short-read Illumina RNA-seq data from a human induced pluripotent stem cell- derived retinal pigmented epithelium (iPSC-RPE) cell line. We use long-read data to identify simple metrics for assessing de novo transcriptome construction and optimize a short-read based de novo transcriptome construction pipeline. We apply this this pipeline to construct transcriptomes for 340 short-read RNA-seq samples originating from healthy adult and fetal human retina, cornea, and RPE. We identify hundreds of novel gene isoforms and examine their significance in the context of ocular development and disease.

genomics