Search bioRxiv⌕ Search

Biology subjects

Vemuri, P.

Publications and source records attributed to Vemuri, P..

2 recordsLinked to original sources

Single-cell transcriptomics for the 99.9% of species without reference genomes

Single-cell RNA-seq (scRNA-seq) is a powerful tool for cell type identification but is not readily applicable to organisms without well-annotated reference genomes. Of the approximately 10 million animal species predicted to exist on Earth, >99.9% do not have any submitted genome assembly. To enable scRNA-seq for the vast majority of animals on the planet, here we introduce the concept of "k-mer homology," combining biochemical synonyms in degenerate protein alphabets with uniform data subsampling via MinHash into a pipeline called Kmermaid. Implementing this pipeline enables direct detection of similar cell types across species from transcriptomic data without the need for a reference genome. Underpinning Kmermaid is the tool Orpheum, a memory-efficient method for extracting high-confidence protein-coding sequences from RNA-seq data. After validating Kmermaid using datasets from human and mouse lung, we applied Kmermaid to the Chinese horseshoe bat (Rhinolophus sinicus), where we propagated cellular compartment labels at high fidelity. Our pipeline provides a high-throughput tool that enables analyses of transcriptomic data across divergent species transcriptomes in a genome- and gene annotation-agnostic manner. Thus, the combination of Kmermaid and Orpheum identifies cell type-specific sequences that may be missing from genome annotations and empowers molecular cellular phenotyping for novel model organisms and species.

bioinformatics↗

Fully Bayesian longitudinal unsupervised learning for the assessment and visualization of AD heterogeneity and progression

Tau pathology and regional brain atrophy are the closest correlate of cognitive decline in Alzheimers disease (AD). Understanding heterogeneity and longitudinal progression of brain atrophy during the disease course will play a key role in understanding AD pathogenesis. We propose a framework for longitudinal clustering that: 1) incorporates whole brain data, 2) leverages unequal visits per individual, 3) compares clusters with a control group, 4) allows to study confounding effects, 5) provides clusters visualization, 6) measures clustering uncertainty, all these simultaneously. We used amyloid-{beta} positive AD and negative healthy subjects, three longitudinal sMRI scans (cortical thickness and subcortical volume) over two years. We found 3 distinct longitudinal AD brain atrophy patterns: a typical diffuse pattern (n=34, 47.2%), and 2 atypical patterns: Minimal atrophy (n=23 31.9%) and Hippocampal sparing (n=9, 12.5%). We also identified outliers (n=3, 4.2%) and observations with uncertain classification (n=3, 4.2%). The clusters differed not only in regional distributions of atrophy at baseline, but also longitudinal atrophy progression, age at AD onset, and cognitive decline. A framework for the longitudinal assessment of variability in cohorts with several neuroimaging measures was successfully developed. We believe this framework may aid in disentangling distinct subtypes of AD from disease staging.

neuroscience↗