Search bioRxiv⌕ Search

Biology subjects

Haregu, S.

Publications and source records attributed to Haregu, S..

2 recordsLinked to original sources

Biological recognition of mirror-image glycans

Recent synthesis of essential enzymes, such as DNA and RNA polymerases with opposite chirality, has boosted the feasibility of creating mirror-image life. Such life, if ever produced, will undoubtedly be coated by a dense display of glycans (glycoproteins, glycolipids, polysaccharides) built from enantiomers of common monosaccharides. Recognition of mirror-image glycans by extant glycan-binding proteins (GBPs) may be critical for colonization by or immune response to mirror life organisms. We evaluated recognition of enantiomers of common glycans by a diverse set of purified GBPs (plant and human derived), antibodies (including IgM from human plasma), mammalian cells (including immune cells), and organs in live animals. We found that GBP binding to enantiomers of naturally prevalent glycans is widespread. Notably, L-glucose and L-galactose interact with fucose-binding lectins, including DC-SIGN, a C-type lectin expressed on immune cells. These interactions can be inhibited by soluble "natural" glycan ligands and enantiomeric ones confirming specificity. Binding of L-glycans to diverse immune cell repertoires revealed preferences for specific glycan enantiomers. IgM antibodies from human serum showed donor-specific recognition of L-glycans. We propose that the recognition of L-glycans by extant GBPs arises from their co-evolution over millennia with the L-glycans that are present in the glycocalyx of many microorganisms.

biochemistry↗

Development of Deep-Learning Models that Predict Quantitative Protein-Ligand Interac-tions in Glycobiology as a part of a Capstone Course

Glycans coat the surface of all cells, and every glycan is recognised by specific glycan-binding proteins (GBPs). There are no general tools that can accurately estimate the binding strength between glycan and GBP from the amino acid sequence of the GBP and the molecular structure of the glycan, represented as SMILES string. We describe models for predicting such binding strengths developed as a part of a Capstone Course at the University of Alberta. The models are trained on a dataset that combines BindingDB, a published database of small-molecule protein interactions, and data from glycan arrays measured by Consortium of Functional Glycomics (CFG). In this hybrid dataset of protein-ligand interactions the ligands are both glycans from CFG and small molecules from BindingDB; similarly, proteins include GBP and proteins from BindingDB. Three models are presented (i) ProMax which fuses ESM-2, MolFormer, and MolCLR features; (ii) APEX which constrains learning to a predetermined form, a physical model of binding; (iii) UltraMax adds inter-atomic distances for the ligands. To address the datasets severe long-tail distribution, the models employ tail-aware losses for rare high-binding instances. Trained and evaluated on approximately one million protein-ligand pairs using hold-out splits for unseen molecules, the three models provide a unified framework for quantitative glycan-protein binding prediction. We observed that learning glycan-protein binding is harder than the similar task of learning small-molecule-protein interactions. Simple mirror-inversion tests led us to postulate that insufficient use of chiral features is an important source of difficulty in learning these interactions.

bioinformatics↗