Search bioRxiv⌕ Search

Biology subjects

Satyajith, A.

Publications and source records attributed to Satyajith, A..

3 recordsLinked to original sources

Development of Deep-Learning Models that Predict Quantitative Protein-Ligand Interac-tions in Glycobiology as a part of a Capstone Course

Glycans coat the surface of all cells, and every glycan is recognised by specific glycan-binding proteins (GBPs). There are no general tools that can accurately estimate the binding strength between glycan and GBP from the amino acid sequence of the GBP and the molecular structure of the glycan, represented as SMILES string. We describe models for predicting such binding strengths developed as a part of a Capstone Course at the University of Alberta. The models are trained on a dataset that combines BindingDB, a published database of small-molecule protein interactions, and data from glycan arrays measured by Consortium of Functional Glycomics (CFG). In this hybrid dataset of protein-ligand interactions the ligands are both glycans from CFG and small molecules from BindingDB; similarly, proteins include GBP and proteins from BindingDB. Three models are presented (i) ProMax which fuses ESM-2, MolFormer, and MolCLR features; (ii) APEX which constrains learning to a predetermined form, a physical model of binding; (iii) UltraMax adds inter-atomic distances for the ligands. To address the datasets severe long-tail distribution, the models employ tail-aware losses for rare high-binding instances. Trained and evaluated on approximately one million protein-ligand pairs using hold-out splits for unseen molecules, the three models provide a unified framework for quantitative glycan-protein binding prediction. We observed that learning glycan-protein binding is harder than the similar task of learning small-molecule-protein interactions. Simple mirror-inversion tests led us to postulate that insufficient use of chiral features is an important source of difficulty in learning these interactions.

bioinformatics↗

Pareto-optimal synthesis of multiple glycans in Golgi compartments

Tree-shaped sugar chains called glycans are covalently attached to proteins in the plasma membrane of eukaryotic cells. They engage in multiple functions such as adhesion, signaling, etc. at the cell surface, so their correct manufacture is of vital importance. Glycans are assembled step by step through enzyme-catalyzed monomer addition reactions as they transit through compartments of the Golgi apparatus. The constraint of fixed residence times across Golgi compartments creates tradeoffs: a short residence time does not give complex glycans enough time for reactions to run to completion, while a long residence time might execute undesirable reactions, which inevitably occur due to enzyme promiscuity. It is not clear how the Golgi reconciles this trade-off. To study this, we devise a from-first-principles model of glycan manufacture with a focus on maximizing glycan yield. Pareto-optimal solutions are a class of solutions that reconcile trade-offs in simultaneously maximizing the yield of multiple glycans. We explore Pareto-optimal solutions for multiple glycan manufacture, and show that small changes in residence times or relative glycan yields can change the optimal enzyme distribution. Together, these results establish Pareto optimality as a unifying framework for interpreting Golgi compartment organization, and for controlling glycan outputs in industrial settings such as manufacture of biotherapeutics, many of which are glycosylated.

systems biology↗

Cell-surface glycans are quantitative reporters of Golgi dysfunction in single cells

Complex sugar polymers known as glycans contribute to the high molecular diversity of the eukaryotic cell surface. The types and levels of glycans on one cell can be sensed by other cells using carbohydrate-specific binding proteins such as lectins, enabling glycans to regulate cell-cell interactions. Glycans are covalently assembled onto proteins and lipids as they traverse the secretory pathway, a tightly regulated process known as glycosylation and carried out by Golgi-resident enzymes. Errors in glycosylation due to the dysfunctional trafficking of enzymes and substrates in the Golgi are implicated in human diseases. Here we ask how much information about Golgi dysfunction is encoded by surface glycans in single cells. This task is challenging due to high cell-to-cell variability in glycan levels. We exploited the loss-of-adhesion driven disorganization of the Golgi in mouse fibroblasts to generate a highly reproducible gradient of Golgi morphology phenotypes, by titrating the Arf1 inhibitor Brefeldin A (BFA). We measured the resulting distribution of cell-surface glycans in single cells using two fluorescently tagged lectin probes (ConA and WGA). A mathematical model of intracellular traffic, parameterized against measurements of Golgi fragmentation and endocytosis, quantitatively explains cell-surface lectin levels across time and BFA concentrations. We used this model to construct an optimal Bayesian decoder and showed that singlecell lectin measurements predict Golgi phenotypes with accuracy far greater than chance. By combining signals from two lectins we further improved prediction accuracy and speed. Such multi-lectin information may be exploited during natural cell-cell communication, and in the development of single-cell diagnostics.

cell biology↗