Search bioRxiv⌕ Search

Biology subjects

Dhodapkar, R. M.

Publications and source records attributed to Dhodapkar, R. M..

5 recordsLinked to original sources

Cell2Sentence: Teaching Large Language Models the Language of Biology

We introduce Cell2Sentence (C2S), a novel method to directly adapt large language models to a biological context, specifically single-cell transcriptomics. By transforming gene expression data into "cell sentences," C2S bridges the gap between natural language processing and biology. We demonstrate cell sentences enable the fine-tuning of language models for diverse tasks in biology, including cell generation, complex cell-type annotation, and direct data-driven text generation. Our experiments reveal that GPT-2, when fine-tuned with C2S, can generate biologically valid cells based on cell type inputs, and accurately predict cell types from cell sentences. This illustrates that language models, through C2S fine-tuning, can acquire a significant understanding of single-cell biology while maintaining robust text generation capabilities. C2S offers a flexible, accessible framework to integrate natural language processing with transcriptomics, utilizing existing models and libraries for a wide range of biological applications.

bioinformatics↗

BrainLM: A foundation model for brain activity recordings

AO_SCPLOWBSTRACTC_SCPLOWWe introduce the Brain Language Model (BrainLM), a foundation model for brain activity dynamics trained on 6,700 hours of fMRI recordings. Utilizing self-supervised masked-prediction training, BrainLM demonstrates proficiency in both fine-tuning and zero-shot inference tasks. Fine-tuning allows for the accurate prediction of clinical variables like age, anxiety, and PTSD as well as forecasting of future brain states. Critically, the model generalizes well to entirely new external cohorts not seen during training. In zero-shot inference mode, BrainLM can identify intrinsic functional networks directly from raw fMRI data without any network-based supervision during training. The model also generates interpretable latent representations that reveal relationships between brain activity patterns and cognitive states. Overall, BrainLM offers a versatile and interpretable framework for elucidating the complex spatiotemporal dynamics of human brain activity. It serves as a powerful "lens" through which massive repositories of fMRI data can be analyzed in new ways, enabling more effective interpretation and utilization at scale. The work demonstrates the potential of foundation models to advance computational neuroscience research.

neuroscience↗

A deep generative model of the SARS-CoV-2 spike protein predicts future variants

AO_SCPLOWBSTRACTC_SCPLOWSARS-CoV-2 has demonstrated a robust ability to adapt in response to environmental pressures--increasing viral transmission and evading immune surveillance by mutating its molecular machinery. While viral sequencing has allowed for the early detection of emerging variants, methods to predict mutations before they occur remain limited. This work presents SpikeGPT2, a deep generative model based on ProtGPT2 and fine-tuned on SARS-CoV-2 spike (S) protein sequences deposited in the NIH Data Hub before May 2021. SpikeGPT2 achieved 88.8% next-residue prediction accuracy and successfully predicted amino acid substitutions found only in a held-out set of spike sequences deposited on or after May 2021, to which SpikeGPT2 was never exposed. When compared to several other methods, SpikeGPT2 achieved the best performance in predicting such future mutations. SpikeGPT2 also predicted several novel variants not present in the NIH SARS-CoV-2 Data Hub. A binding affinity analysis of all 54 generated substitutions identified 5 (N439A, N440G, K458T, L492I, and N501Y) as predicted to simultaneously increase S/ACE2 affinity, and decrease S/tixagevimab+cilgavimab affinity. Of these, N501Y has already been well-described to increase transmissibility of SARS-CoV-2. These findings indicate that SpikeGPT2 and other similar models may be employed to identify high-risk future variants before viral spread has occurred.

bioinformatics↗

Representing cells as sentences enables natural-language processing for single-cell transcriptomics

AO_SCPLOWBSTRACTC_SCPLOWGene expression matrices commonly used in single-cell transcriptomics, cannot be directly analyzed with tools developed for natural languages. By restructuring these matrices as abundance-ordered sequences of genes, we generate cell sentences: rank-normalized, positionally encoded expression data. We show that these cell sentences can be analyzed using existing tools from natural language processing to unify cell and gene representations across species.

bioinformatics↗

Causal identification of single-cell experimental perturbation effects with CINEMA-OT

Recent advancements in single-cell technologies allow characterization of experimental perturbations at single-cell resolution. While methods have been developed to analyze such experiments, the application of a strict causal framework has not yet been explored for the inference of treatment effects at the single-cell level. In this work, we present a causal inference based approach to single-cell perturbation analysis, termed CINEMA-OT (Causal INdependent Effect Module Attribution + Optimal Transport). CINEMA-OT separates confounding sources of variation from perturbation effects to obtain an optimal transport matching that reflects counterfactual cell pairs. These cell pairs represent causal perturbation responses permitting a number of novel analyses, such as individual treatment effect analysis, response clustering, attribution analysis, and synergy analysis. We benchmark CINEMA-OT on an array of treatment effect estimation tasks for several simulated and real datasets and show that it outperforms other single-cell perturbation analysis methods. Finally, we perform CINEMA-OT analysis of two newly-generated datasets: (1) rhinovirus and cigarette smoke-exposed airway organoids, and (2) combinatorial cytokine stimulation of immune cells. In these experiments, CINEMA-OT reveals potential mechanisms by which cigarette smoke exposure dulls the airway antiviral response, as well as the logic that governs chemokine secretion and peripheral immune cell recruitment.

bioinformatics↗