Search bioRxiv⌕ Search

Biology subjects

Karaletsos, T.

Publications and source records attributed to Karaletsos, T..

7 recordsLinked to original sources

A Cross-Species Generative Cell Atlas Across 1.5 Billion Years ofEvolution: The TranscriptFormer Single-cell Model

Single-cell transcriptomics has revolutionized our understanding of cellular diversity, yet our understanding of the transcriptional programs across the tree of life remains limited. Here we present TranscriptFormer, a family of generative foundation models trained on up to 112 million cells spanning 1.53 billion years of evolution across 12 species. By jointly modeling gene identities and expression levels using a novel generative architecture, TranscriptFormer encodes multi-scale biological structure, functioning as a queryable virtual cell atlas. We demonstrate state-of-the-art performance on both in-distribution and out-of-distribution cell type classification, with robust performance even for species separated by over 685 million years of evolution. TranscriptFormer can also perform zero-shot disease state identification in human cells and accurately transfers cell state annotations across species boundaries. As a generative model, TranscriptFormer can be prompted to predict cell type-specific transcription factors and gene-gene interactions that align with independent experimental observations. Developmental trajectories, phylogenetic relationships and cellular hierarchies emerge naturally in TranscriptFormers representations without any explicit training on these annotations. This work establishes a powerful framework for quantitative single-cell analysis, and comparative cellular biology, thus demonstrating that universal principles of cellular organization can be learned and predicted across the tree of life.

systems biology↗

SubCell: Vision foundation models for microscopycapture single-cell biology

Cell morphology and subcellular protein organization provide important insights into cellular function and behavior. These cellular features can be studied using large-scale fluorescence microscopy, and machine learning has become a powerful tool to interpret the resulting images for biological insights. Here, we introduce SubCell, a deep learning model for fluorescence microscopy designed to accurately capture cellular morphology, protein localization, cellular forganization, and biological function beyond what humans can readily perceive. SubCell was trained on the proteome-wide image collection from the Human Protein Atlas with a novel proteome-aware learning objective. SubCell outperforms state-of-the-art methods across a variety of tasks relevant to single-cell biology and generalizes to other fluorescence microscopy datasets without any fine-tuning. Additionally, we construct the first proteome-wide hierarchical map of proteome organization that is directly learned from image data. This vision-based multiscale cell map defines cellular subsystems down to protein complex resolution, reveals proteins with similar functions, and distinguishes dynamic and stable behaviors within cellular compartments. Finally, combining SubCell with a protein sequence model enables a rich multimodal approach to capture gene function better than either vision-only or sequence-only models alone. In conclusion, SubCell creates deep, image-driven representations of cellular architecture that are applicable across diverse biological contexts and datasets.

cell biology↗

scGenePT: Is language all you need for modeling single-cell perturbations?

Modeling single-cell perturbations is a crucial task in the field of single-cell biology. Predicting the effect of up or down gene regulation or drug treatment on the gene expression profile of a cell can open avenues in understanding biological mechanisms and potentially treating disease. Most foundation models for single-cell biology learn from scRNA-seq counts, using experimental data as a modality to generate gene representations. Similarly, the scientific literature holds a plethora of information that can be used in generating gene representations using a different modality - language - as the basis. In this work, we study the effect of using both language and experimental data in modeling genes for perturbation prediction. We show that textual representations of genes provide additive and complementary value to gene representations learned from experimental data alone in predicting perturbation outcomes for single-cell data. We find that textual representations alone are not as powerful as biologically learned gene representations, but can serve as useful prior information. We show that different types of scientific knowledge represented as language induce different types of prior knowledge. For example, in the datasets we study, subcellular location helps the most for predicting the effect of single-gene perturbations, and protein information helps the most for modeling perturbation effects of interactions of combinations of genes. We validate our findings by extending the popular scGPT model, a foundation model trained on scRNA-seq counts, to incorporate language embeddings at the gene level. We start with NCBI gene card and UniProt protein summaries from the genePT approach and add gene function annotations from the Gene Ontology (GO). We name our model "scGenePT", representing the combination of ideas from these two models. Our work sheds light on the value of integrating multiple sources of knowledge in modeling single-cell data, highlighting the effect of language in enhancing biological representations learned from experimental data.

bioinformatics↗

Deep Learning Analysis on Images of iPSC-derived Motor Neurons Carrying fALS-genetics Reveals Disease-Relevant Phenotypes

Amyotrophic lateral sclerosis (ALS) is a devastating condition with very limited treatment options. It is a heterogeneous disease with complex genetics and unclear etiology, making the discovery of disease-modifying interventions very challenging. To discover novel mechanisms underlying ALS, we leverage a unique platform that combines isogenic, induced pluripotent stem cell (iPSC)-derived models of disease-causing mutations with rich phenotyping via high-content imaging and deep learning models. We introduced eight mutations that cause familial ALS (fALS) into multiple donor iPSC lines, and differentiated them into motor neurons to create multiple isogenic pairs of healthy (wild-type) and sick (mutant) motor neurons. We collected extensive high-content imaging data and used machine learning (ML) to process the images, segment the cells, and learn phenotypes. Self-supervised ML was used to create a concise embedding that captured significant, ALS-relevant biological information in these images. We demonstrate that ML models trained on core cell morphology alone can accurately predict TDP-43 mislocalization, a known phenotypic feature related to ALS. In addition, we were able to impute RNA expression from these image embeddings, in a way that elucidates molecular differences between mutants and wild-type cells. Finally, predictors leveraging these embeddings are able to distinguish between mutant and wild-type both within and across donors, defining cellular, ML-derived disease models for diverse fALS mutations. These disease models are the foundation for a novel screening approach to discover disease-modifying targets for familial ALS.

neuroscience↗

EmbedGEM: A framework to evaluate the utility of embeddings for genetic discovery

Machine learning (ML)-derived embeddings are a compressed representation of high content data modalities. Embeddings can capture detailed information about disease states and have been qualitatively shown to be useful in genetic discovery. Despite their promise, embeddings have a major limitation: it is unclear if genetic variants associated with embeddings are relevant to the disease or trait of interest. In this work we describe EmbedGEM (Embedding Genetic Evaluation Methods), a framework to systematically evaluate the utility of embeddings in genetic discovery. EmbedGEM focuses on comparing embeddings along two axes: heritability and disease relevance. As measures of heritability, we consider the number of genome-wide significant associations and the mean{chi} 2 statistic at significant loci. For disease relevance, we compute polygenic risk scores for each embedding principal component, then evaluate their association with high-confidence disease or trait labels in a held-out evaluation patient set. While our development of EmbedGEM is motivated by embeddings, the approach is generally applicable to multivariate traits, and can readily be extended to accommodate additional metrics along the evaluation axes. We demonstrate EmbedGEMs utility by evaluating embeddings and multivariate traits in two separate datasets: i) a synthetic dataset simulated to demonstrate the ability of the framework to correctly rank traits based on their heritability and disease relevance, and ii) a real data from the UK Biobank including metabolic and liver-related traits. Importantly, we show that greater disease relevance does not automatically follow from greater heritability.

bioinformatics↗

Pitfalls in performing genome-wide association studies on ratio traits

Genome-wide association studies (GWAS) are often performed on ratios composed of a numerator trait divided by a denominator trait. Examples include body mass index (BMI) and the waist-to-hip ratio, among many others. Explicitly or implicitly, the goal of forming the ratio is typically to adjust for an association between the numerator and denominator. While forming ratios may be clinically expedient, there are several important issues with performing GWAS on ratios. Forming a ratio does not "adjust" for the denominator in the sense of conditioning on it, and it is unclear whether associations with ratios are attributable to the numerator, the denominator, or both. Here we demonstrate that associations arising in ratio GWAS can be entirely denominator-driven, implying that at least some associations uncovered by ratio GWAS may be due solely to a putative adjustment variable. In a survey of 10 common ratio traits, we find that the ratio model disagrees with the adjusted model (performing GWAS on the numerator while conditioning on the denominator) at around 1/3 of loci. Using BMI as an example, we show that variants detected by only the ratio model are more strongly associated with the denominator (height), while variants detected by only the adjusted model are more strongly associated with the numerator (weight). Although the adjusted model provides effect sizes with a clearer interpretation, it is susceptible to collider bias. We propose and validate a simple method of correcting for the genetic component of collider bias via leave-one-chromosome-out polygenic scoring.

genetics↗

An allelic series rare variant association test for candidate gene discovery

Allelic series are of candidate therapeutic interest due to the existence of a dose-response relationship between the functionality of a gene and the degree or severity of a phenotype. We define an allelic series as a gene in which increasingly deleterious mutations lead to increasingly large phenotypic effects, and develop a gene-based rare variant association test specifically targeted for the identification of allelic series. Building on the well-known burden and sequence kernel association (SKAT) tests, we specify a variety of association models, covering different genetic architectures, and integrate these into a COding-variant Allelic Series Test (COAST). Through extensive simulations, we confirm that COAST maintains the type I error and improves power when the pattern of coding-variant effect sizes increases monotonically with mutational severity. We applied COAST to identify allelic series for 4 circulating lipid traits and 5 cell count traits among 145,735 subjects with available whole exome sequencing data from the UK Biobank. Compared with optimal SKAT (SKAT-O), COAST identified 29% more Bonferroni significant associations with circulating lipid traits, on average, and 82% more with cell count traits. All of the gene-trait associations identified by COAST have corroborating evidence either from rare-variant associations in the full cohort (Genebass, N = 400K), or from common variant associations in the GWAS catalog. In addition to detecting many gene-trait associations present in Genebass using only a fraction (36.9%) of the sample, COAST detects associations, such as ANGPTL4 with triglycerides, that are absent from Genebass but which have clear common variant support.

genetics↗