Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.06.04.730134

Vibe Coding Specificity Foundation Models

Abstract

Molecular recognition -- the determination of which agent binds which target -- governs adaptive immunity, gene regulation, signal transduction, RNA silencing, enzyme catalysis, and the selectivity of therapeutics. Determining binding specificity remains dependent on experimental screening or domain-specific computational tools that do not generalize across binding modalities. Transformer softmax attention is mathematically identical to the Boltzmann distribution governing molecular binding1. This identity, together with five conditions of molecular recognition systems, prescribes a single neural network architecture for cross-modal binding prediction: dual sequence encoders, symmetric contrastive learning, and a learned physical temperature2. A Specificity Foundation Model (SFM) is an instance of this physics-derived, sequence-to-sequence architecture that maps any agent-target sequence pair to a binding compatibility score, enabling bidirectional retrieval across molecular recognition domains without requiring structural information. The first SFM for antibody-antigen binding demonstrated [~]100,000-fold greater data efficiency than comparable vision-language models3. Here we report six SFMs across six molecular recognition domains -- transcription factor-DNA, enzyme-substrate, peptide-MHC, CRISPR gRNA-off-target genomic DNA, microRNA-mRNA target, and small molecule drug-target protein -- using the identical architecture without modification and trained using publicly available data only. Evaluated by cross-modal retrieval from pools of 512 candidates (random baseline 0.2%), in-distribution R@1 ranges from 27.7% to 98.0% across the six domains. mir-SFM retrieves miRNA targets at 98.0% R@1, including the [~]80% of validated interactions that seed-matching tools cannot find. mhcSFM achieves 95.4% R@1 on held-out rare HLA alleles absent from training. Applying crisprSFM to CRISPR off-target prediction improves precision to 94.0% compared to 33.2% from Hamming distance alone. All six SFMs were built by a domain expert with no programming experience using vibe coding -- natural-language-directed AI coding agents -- with numerical claims independently verified by an orthogonal AI auditor. These results establish SFMs as a physics-derived, sequence-native class of model that augments experimental and computational workflows across molecular recognition domains.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Reddy, S. T.. 2026-06-04. Vibe Coding Specificity Foundation Models. https://doi.org/10.64898/2026.06.04.730134

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Engineering phenotypic heterogeneity for functional organization in microbial populations

Engineering microbial populations to perform complex functions requires programming not only cellular behaviour but also the functional organization through which biological activities are distributed across populations. While phenotypic heterogeneity is often regarded as variability to suppress, it can also serve as a foundation for organizing specialized functions within genetically homogeneous microbial systems. Here, we introduce PROMETEO, a modular genetic-circuit framework that programs population composition through architecture-encoded regulatory and translational asymmetries. Using a library of 27 asymmetric bistable circuits, we demonstrate that circuit architecture reproducibly specifies phenotypic distributions spanning a broad range of population compositions without continuous external induction. Programmed population structures remained stable over serial propagation and were qualitatively conserved across Escherichia coli and Pseudomonas putida. Stochastic and deterministic modelling accurately predicted architecture-dependent population compositions and hysteresis regimes, providing a quantitative framework for rational design. We further show that programmable population composition supports multiple modes of functional organization, including stable parallel specialization, inducible temporal redistribution of cellular states, and spatial ecological compartmentalization through biofilm-associated and planktonic subpopulations. As a demonstration of these capabilities, architecture-programmed organization enabled distributed Congo Red biotransformation through coordinated reductive and oxidative activities, achieving up to 98% dye removal in spatially compartmentalized populations. Together, these results establish population composition as a programmable property of genetic circuit design and provide a general strategy for engineering distributed functions within genetically homogeneous microbial populations.

synthetic biology↗

Carboxysome-Inspired Protein Coacervates for Light-Driven CO2 Reduction and H2 Evolution

Efficient catalysis often requires high local concentrations of reactants and catalysts, which cells achieve through compartmentalization within organelles, such as carboxysomes, that increase the efficiency of bacterial carbon fixation. Here, we engineered a photocatalytic reaction compartment that concentrated an artificial metalloenzyme, carbon dioxide, and a photosensitizer by liquid-liquid phase separation triggered by a cationic polypeptide, deca(L-arginine) (R10). At low R10 concentrations, CoPPIX binding increases the alpha-helical structure of the otherwise disordered protein, supercharged cytochrome b5622(-22). At higher concentrations, electrostatic complexation produces spherical droplets that enrich the protein and cobalt cofactor and recruit the photosensitizer [Ru(bpy)3]2+. Under illumination, coacervation increased hydrogen evolution 1.9-fold and CO formation from carbon dioxide; 1.3-fold relative to the corresponding solution-phase protein system. Co-encapsulation of carbonic anhydrase changed the product distribution specifically in the condensed phase: CO production increased 3.2-fold, hydrogen evolution decreased from 1.51 to 0.70 mol, and CO selectivity among the detected two-electron products rose from 33% to 77%. These results demonstrate that bioinspired coacervates can stabilize reactive intermediates, enrich local substrate concentrations, and integrate multiple catalytic functions, providing a generalizable framework for programmable, light-driven synthetic organelles.

synthetic biology↗

Biocatalytic Production of Galantamine in Yeast through Cytochrome P450 Optimization

Galantamine is a pharmaceutically relevant Amaryllidaceae alkaloid used for the treatment of Alzheimer's disease. Its structural complexity and low abundance in plants motivate the development of alternative manufacturing routes. Here, we established the first engineered yeast platform for the biocatalytic production of galantamine from 4OMe-norbelladine. Heterologous expression of the downstream pathway enzymes NtCYP96T6, NtNMT1, and NtAKR1 from Narcissus cv. Tete-a-Tete in Saccharomyces cerevisiae was complemented by optimization of cultivation temperature, carbon source, and medium pH enabling the first demonstration of galantamine production in yeast. Systematic screening of cytochrome P450 reductase and cytochrome b partners to boost NtCYP96T6 activity led to the identification of a novel reductase mined from the Narcissus pseudonarcissus transcriptome, NpCPR, which supported the highest pathway flux and yielded 7.0 {+/-} 0.6 mg/L galantamine, corresponding to a 7.6 % molar yield from 250 M 4OMe-norbelladine and an approximately 173-fold improvement over the parental strain. We further exploited this yeast cell factory for the precursor-directed biosynthesis of 7F-galantamine, highlighting the potential of pathway enzyme promiscuity to access new-to-nature GAL analogues that may be challenging to produce through conventional chemical synthesis. Together, this work establishes a foundation for microbial galantamine production and biosynthetic diversification of its pharmaceutically relevant scaffold.

synthetic biology↗