Search bioRxiv⌕ Search

Biology subjects

Taraman, S.

Publications and source records attributed to Taraman, S..

3 recordsLinked to original sources

A World Model of Molecular Organization Detects Cryptic Pockets from Apo Structure

Cryptic pockets are druggable sites that are absent in a protein's resting structure and form only upon backbone rearrangement, posing a challenge for detection from static apo structures. Existing methods rely on generating open conformations through sampling or complex prediction, limiting applicability. Here we present a novel detector that reads cryptic pockets directly from a structural world-model latent representation of a single apo structure, without conformational sampling or external pocket finders. Evaluated on CryptoBench and CryptoBank datasets, the method localizes cryptic sites with top-1 accuracy as high as 0.848 and top-5 accuracy exceeding 0.93, and successfully recovers an allosteric site on held-out WRN helicase structures. The detector complements existing approaches and improves with training data scale. These findings suggest that structural world-model latents encode conformational flexibility, enabling effective cryptic pocket identification from apo structures alone, facilitating drug discovery on challenging targets. Benchmarked against the co-folding engine OpenDDE, the detector finds pockets directly rather than by first predicting a bound complex: on 190 targets held out of both training sets it recovers 119 sites OpenDDE misses, against 2 in the other direction, and at top-5 co-folding's recovered sites are a subset of ours. It runs on any structure from the apo coordinates alone, including the mmCIF-only entries all recent depositions carry, and its accuracy keeps climbing as the training corpus grows. The intended use is prospective cryptic-site nomination on the targets that sequence and static structure leave without a starting point.

molecular biology↗

Autonomous AI-Driven Nanoscale Spatial Mapping Reveals Novel Targets and Ternary Architectures in 5xFAD Alzheimer's Disease Model

Alzheimer's disease (AD) is characterized by the deposition of amyloid-beta (A{beta}) and microtubule-associated protein tau (MAPT) neurofibrillary tangles in the brain; however, the molecular mechanisms underlying associated synaptic dysfunction remain unclear. Eratos' AI for Spatial Computing and Embedded Neurotherapeutic Discovery (ASCENDTM) engine was employed to analyze multiplexed expansion revealing (multiExR) data from the 5xFAD and wild-type mouse somatosensory cortex, enabling quantitative mapping of A{beta}, RIM1, and GluA2 nanodomains. ASCENDTM identified significant and previously unreported protein associations, including nanoscale colocalization of A{beta} with postsynaptic AMPA-receptor subunit GluA2 and amyloid-bridged A{beta}-GluA2-RIM1 ternary assemblies. These spatial signatures indicate complex synaptic disruption, defined by distinct morphological and density profiles associated with neurobiological and neuroinflammatory pathology. Integrating high-dimensional spatial computing with cross-modal literature synthesis advances traditional microscopy analysis toward autonomous, AI-driven scientific discovery. Identifying novel, druggable interfaces within native tissue expedites the discovery of precision neurotherapeutics for AD and other complex central nervous system disorders.

systems biology↗

HI-JEPA: A World Model of Molecular Organization Learned from Measured Proximity

Proteins act through the company they keep. Which molecules occupy the same nanoscale neighborhood in intact tissue determines what can physically interact, and disease rearranges those neighborhoods before it changes anything a sequence records. That quantity (measured proximity between molecular species in unperturbed tissue) has never been acquired broadly enough to train on. Published colocalization arrives study by study and never accumulates into a graph. The measurement has to be made rather than collected. We built ASCEND, a spatial computing platform that measures pairwise molecular proximity from expansion microscopy at molecular resolution in intact tissue, and applied it to 164 proteins across 37 imaged regions in five studies, spanning cultured neurons, isolated synapses and mouse cortex in disease and control. HI-JEPA is a representation trained on those measurements. Each protein is one embedding, trained to predict the embeddings of its measured neighbors in latent space; it never reconstructs its input and generates no negatives. A set of proteins measured in one neighborhood forms a configuration, which is the object the model perturbs and plans over. The representation performs operations a sequence model cannot. It names a protein from the bare geometry of a microscopy point cloud, matched against 234,048 deposited structures, at top-1 accuracy 0.748 against a chance rate of 1.0 x 10-5. It predicts physical interaction between sequence-dissimilar proteins that were both withheld from training at AUC 0.908, where ESM-C 6B reaches 0.514 against partner-count-matched negatives. It recovers a held-out complex member in the top 100 of 13,447 candidates at recall 0.954, against 0.514 for a ranking built from complex frequency alone. Asked which partners a knockout disrupts, it recovers the experimentally observed ones at recall@100 0.640; asked the same question about a different protein, with the ranking rule and denominators unchanged, it recovers 0.028, so the answer follows the action. Given 5xFAD mouse cortex with no disease label, no reward and no indication that amyloid is relevant, ranking 1,574 measured assemblies by their departure from wild type returns amyloid-{beta} bound to AMPA receptor subunits in nine of the top ten. Planning over the same configurations independently selects the same subunits (GluA2, GluA3, GluA4) and predicts that disrupting the PSD-95 scaffold worsens the configuration, both agreeing in sign with experiments the model never saw. Ablating the measured-proximity channel at training time degrades cross-scale partner recovery from median rank 14 to 68 while leaving navigation and within-scale dynamics intact; ablating the perturbation channel does the reverse. The cross-scale capability therefore comes from the measurement and not from having seen more data. The intended application is target nomination in diseases where sequence and structure supply no starting point. Note: This is a capability report. The architecture, the training procedure and the acquisition protocol are proprietary and are not described. Section 4.2 gives the evaluation protocol behind every number reported.

systems biology↗