Search bioRxiv⌕ Search

Biology subjects

St John, J.

Publications and source records attributed to St John, J..

2 recordsLinked to original sources

Interpretable gene networks from single-cell foundation models reveal conserved neurogenic dysfunction in Parkinson's disease

Interpreting large-scale single-cell transcriptomic data remains a major challenge for understanding disease mechanisms. Recent single-cell foundation models learn rich representations of gene relationships across millions of cells, yet methods for translating these embeddings into biologically interpretable gene networks remain limited. Here we present scGENet, a computational framework that constructs context-specific gene interaction networks from foundation model-derived gene embeddings. By fine-tuning pretrained models on transcriptomic data from human midbrain organoids, scGENet generates transcriptome-scale gene modules that capture biologically meaningful cellular programs. Benchmarking across multiple foundation models demonstrates that networks derived from a fine-tuned scGPT brain model show the highest concordance with curated neuronal pathways, Parkinsons disease (PD) genetic risk loci, and independent patient-derived transcriptional signatures. Applying this framework to human iPSC-derived PD midbrain organoids reveals transcriptional modules associated with neuronal differentiation, synaptic signaling, and cell-cycle regulation. Single-nucleus RNA sequencing further links these programs to altered cellular composition, including reduced dopaminergic neurons, expansion of radial glia-like progenitors, and a dopaminergic neuron subtype expressing SNCA and VGLUT2. Integration with independent human substantia nigra datasets identifies a conserved neurogenic program disrupted across genetic and idiopathic PD. Together, these results establish a generalizable strategy for extracting interpretable gene networks from single-cell foundation models, enabling systematic discovery of disease-relevant molecular programs across diverse tissues and datasets.

neuroscience↗

Designing AI-programmable therapeutics with the EDEN family of foundation models

The ability to interpret, modify, and design DNA has driven many of the most significant advances in modern medicine, from diagnostics, biologics, and vaccines to cell and gene therapies. However, the inherent complexity of biological systems means that most modern medicines are still engineered using bespoke, labor-intensive processes. To address the need for a generalisable and programmable approach to therapeutic design, we introduce the EDEN (environmentally-derived evolutionary network) family of metagenomic foundation models, including a 28 billion parameter model trained on 9.7 trillion nucleotide tokens from BaseData1. This dataset, at the time of training, contained more than 10 billion novel genes from over 1 million new species, and is intentionally enriched for environmental and host-associated metagenomes, phage sequences, and mobile genetic elements, enabling the model to learn from diverse and novel cross-species evolutionary mechanisms and apply them to key challenges in human health. EDEN achieves state-of-the-art performance across a series of predictive and generative genomic and protein benchmarks. To demonstrate the models broad applicability across biology, we evaluate EDENs capacity for programmable therapeutic design by challenging a single architecture to design biological novelty across three distinct therapeutic modalities, disease areas and biological scales: (i) large gene insertion, (ii) antibiotic peptide design, and (iii) microbiome design. First, we demonstrate AI-programmable Gene Insertion (aiPGI), in which EDEN designs de novo large serine recombinases (LSRs) capable of inserting large pieces of DNA at desired target sites in the human genome when prompted only on 30 nucleotides of DNA sequence from the desired target site. In low-N experimental validation, EDEN generated multiple active recombinases for all tested disease-associated genomic loci (ATM, DMD, F9, FANCC, GALC, IDS, P4HA1, PHEX, RYR2, USH2A) and 4 potential safe harbor sites in the human genome. EDEN achieves an overall functional hit rate of 63.2% across diverse DNA prompts when prompted on only 30bp of DNA from outside the training data. 50% of EDEN-generated LSRs were active in human cells, achieving therapeutically relevant levels of CAR insertion in primary human T cells. We also show that EDEN can generate active bridge recombinases when prompted on the associated guide RNA alone, with sequence identities to training and public data as low as 65%. These results pave the way for a new generation of cell and gene therapies by opening the door to rapid, programmable and site-specific integration of large genetic payloads without double-strand breaks. This offers an alternative to the safety, efficiency and payload limitations inherent in viral or nuclease-based editing at thousands of currently intractable human therapeutic targets. Second, we use the same model to generate a focused low-N library of novel antimicrobial peptides where 97% showed activity, with top candidates achieving single-digit micromolar potency against critical-priority multidrug-resistant pathogens. Third, to demonstrate that EDEN captures inter-genomic features, we design a gigabase-scale microbiome with over 94,000 synthetic metagenomic assemblies, including prophage genomes and correct cross-species metabolic pathway completions. The EDEN-generated synthetic microbiome covers 9,067 species with a biome-specific taxonomic accuracy of 99%. Over 1,500 of the generated species were outside the fine-tuning dataset while retaining the correct microecological properties and biome association, thus significantly expanding genetic and taxonomic diversity. Together, these results establish a new strategic direction for AI-programmable therapeutics, in which a single foundation model architecture designs candidate therapeutics across diverse modalities and disease areas. This suggests that the combination of billions of years of evolutionary data with specific therapeutic records offers a clear, scaling-driven path to making therapeutic design a predictable engineering discipline. O_FIG O_LINKSMALLFIG WIDTH=141 HEIGHT=200 SRC="FIGDIR/small/699009v2_ufig1.gif" ALT="Figure 1"> View larger version (59K): org.highwire.dtl.DTLVardef@68c20borg.highwire.dtl.DTLVardef@19b8e8corg.highwire.dtl.DTLVardef@1ab9362org.highwire.dtl.DTLVardef@1592cb7_HPS_FORMAT_FIGEXP M_FIG C_FIG

genomics↗