Search bioRxiv⌕ Search

Biology subjects

Huynh, G.

Publications and source records attributed to Huynh, G..

2 recordsLinked to original sources

PULSAR: a Foundation Model for Multi-scale and Multicellular Biology

Biology emerges from interactions across physical scales, where molecular interactions drive cellular states, which in turn orchestrate multicellular tissue functions that collectively define health and disease. However, current computational models are often constrained to single scales in isolation, failing to integrate the biology that emerges from lower to higher levels [1]. Here we present PULSAR (Patient Understanding Leveraging Single-cell universAl Representation), a multi-scale and multicellular foundation model architecture that explicitly enables information flow from genes to cells to multicellular systems. Applied to the human peripheral immune system, PULSAR extracts a unified donor representation that supports rapid disease classification, biomarker prediction, and forecasting of future clinical events, such as Rheumatoid arthritis onset. As a generative model, PULSAR enables the simulation of cytokine perturbation response across physical resolutions, while its interpretability reveals the key cell types driving disease. Overall, PULSAR opens new avenues for precision medicine by enabling computational reasoning that connects molecular biology to clinical phenotypes.

bioinformatics↗

TabVI: Leveraging Lightweight Transformer Architectures to Learn Biologically Meaningful Cellular Representations

Transformer-based foundation models are changing the landscape of natural language processing (NLP), computer vision, and audio, achieving human-level performance across a variety of tasks. Extending these models to single-cell genomics holds significant potential for revealing the cellular and molecular perturbations associated with disease. However, unlike the sequential structure of language, the functional organization of genes is hierarchical and modular. This fundamental difference necessitates the development of meaningful feature selection strategies to adapt NLP transformer architectures effectively. In contrast to many large-scale foundation models, probabilistic models have shown success in learning complex cellular representations from single-cell datasets. In this work, we present TabVI, a probabilistic deep generative model that leverages tabular transformer architectures to improve latent embedding learning. We validate TabVIs performance in cell type annotation and integration benchmarks. We demonstrate that TabVI improves performance across down-stream tasks and is robust to scaling dataset sizes, producing interpretable, sample-specific feature attention masks. TabVI is a lightweight, scientifically-meaningful, transformer architecture for single-cell analysis that excels where large scale foundation models are less effective.

bioinformatics↗