Search bioRxiv⌕ Search

Biology subjects

Arts, M.

Publications and source records attributed to Arts, M..

4 recordsLinked to original sources

EpiRanha: Hunting for Epitope Similarity with a Structure- and Residue-Aware Graph Neural Network

AO_SCPLOWBSTRACTC_SCPLOWPrecise epitope recognition underpins the efficacy and safety of therapeutic antibodies, yet existing approaches to epitope similarity scoring rely largely on sequence identity or rigid structural superposition, limiting their ability to robustly assess cross-reactivity and potential off-target interactions. We introduce EpiRanha, a multimodal framework that integrates residue-level ESM-2 sequence embeddings with an E(n)-equivariant graph neural network operating on three-dimensional protein structure. EpiRanha produces per-residue "fingerprints" that jointly encode sequence-level context and spatial organization, and then applies a beam-search strategy to identify and rank multiple high-confidence epitope candidates across protein surfaces that are similar to a given query epitope. We evaluate EpiRanha against TM-align on nanobody-antigen complexes from SAbDab-nano and a set of AlphaFold-predicted proteins. EpiRanha consistently recovers the query epitope on its cognate antigen, including highly discontiguous conformational epitopes that rigid alignment methods such as TM-align often fail to capture, achieving lower structural loss and fewer false negatives through flexible residue-level mapping. Beyond these self-matches, EpiRanha also identifies biologically plausible epitope-level similarities on other proteins. Overall, EpiRanha advances epitope characterization beyond sequence or geometry alone, enabling more robust off-target risk assessment, informing training-set construction for predictive models, and more selective antibody design.

bioinformatics↗

Effects of protein interface mutations on protein quality and affinity

Accurately modeling antibody-antigen interactions requires distinguishing intrinsic binding affinity ("protein-interaction") from protein biophysical properties ("protein-quality"), including folding, stability, and expression. However, high-throughput mutational measurements commonly used to train and benchmark computational models often conflate these effects, obscuring the true determinants of molecular recognition. Here, we present an experimental and analytical framework to disentangle protein-interaction effects from protein-quality effects in single-domain antibody (VHH)-antigen binding. Using a large-scale deep mutational scanning (DMS) dataset spanning four VHH-antigen complexes, with single and double mutations in both partners, we introduce control binders to quantify protein-quality changes independently of protein-interaction. This enables decomposition of experimentally measured affinity into protein-interaction and protein-quality components at scale. Leveraging the disentangled dataset, we evaluated state-of-the-art structure- and sequence-based models for protein-quality and protein-interaction prediction and show that their performance largely reflects protein-quality rather than protein-interaction effects. Our results highlight a major confounder in current datasets and suggest that accounting for protein-quality will be essential for training next-generation affinity-prediction models. Nomenclature Antibody related termsO_LIPrimary VHH: The VHH of a VHH-antigen complex for which the paratope and the epitope weremutated. C_LIO_LIControl VHH: A second VHH that binds to the same antigen as the primary VHH but has non-overlapping epitope positions and therefore does not bind to any of the mutated antigen positions. C_LI Affinity-related termsO_LIReal Affinity: "The strength of the interaction between two [...] molecules that bind reversibly (interact)" 1. In the context of antibody-antigen binding, it quantifies interactions between active proteins (which are expressed and correctly folded 2 and are therefore functionally and biologically active (see below). It is commonly quantified by the equilibrium dissociation constant, KD. C_LIO_LIObserved affinity ({degrees}KD): The interaction strength experimentally measured between two molecules. Unlike real affinity, this value is confounded by the biophysical properties of the individual binding partners, specifically their folding, stability, and expression levels. Consequently, the observed affinity often differs from the real/intrinsic affinity if a significant fraction of the protein population is inactive 3. NOTE: Unless otherwise specified, {degrees}KD is reported in - log10 space. For example, a {degrees}KD of -9 corresponds to 10-9M or 1nM. C_LIO_LIChange in observed affinity ({Delta}{degrees}KD): The shift in the observed affinity between two proteins upon mutation, reported as the log10-transformed fold change. A value of 1 reflects a 10-fold difference, a value of 2 a 100-fold difference, etc. This aggregate change resolves into two distinct biophysical components 2, 4: O_LIProtein-interaction change: The change in the intrinsic thermodynamic affinity between the two binding partners, each in its active state (i.e., the specific change in interface Gibbs free energy because both enthalpy and entropy are considered). C_LIO_LIProtein-quality change: The change in the fraction of the mutated protein population that is biologically active - meaning it is expressed, correctly folded, and stable 2, 5. O_LIFolding: The process that guides the polypeptide chain toward its native conformation, which is a prerequisite for forming a functional binding site. C_LIO_LIStability: The thermodynamic capacity to maintain the folded structure over time and under physiological conditions. Stability (decrease in Gibbs free energy from the unfolded to the folded state) ensures the binding interface remains intact and prevents competing processes such as aggregation 6. C_LIO_LIExpression: The steady-state abundance of the protein. This is largely dependent on proper folding and stability, as cellular quality control mechanisms degrade proteins that fail to fold or remain stable at functional concentrations. C_LI C_LI C_LIO_LIChange in relative affinity ({Delta}{Delta}{degrees}KD): the difference between the {Delta}{degrees}KD of the primary VHH compared to the control VHH for a given epitope mutation. C_LI Model-related termsO_LIESM-IF1 sc: Single-chain (sc) structure-conditioned inverse folding model (ESM-IF1), using the isolated monomer structure of the mutated protein: either the VHH or the antigen 7. C_LIO_LIESM-IF1 mc: Multi-chain (mc) structure-conditioned model (ESM-IF1), using the full complex structure (both antibody and antigen) 7. C_LIO_LIStability prediction score: Score that represents the predicted change in stability based on a single mutation, normally represented as {Delta}{Delta}G. C_LI

molecular biology↗

The importance of UBQLN2 ubiquitylation for its turnover and localization

UBQLN2 is a member of the UBL-UBA domain protein family that functions as extrinsic substrate receptors for the 26S proteasome. UBQLN2 has been shown to undergo phase separation in vitro. In cells, UBQLN2 forms condensates that may be of importance for tuning protein degradation via the ubiquitin-proteasome system and potentially of relevance for UBQLN2-linked amyotrophic lateral sclerosis (ALS). Here we show that UBQLN2 is ubiquitylated on lysine residues in the N-terminal UBL domain. The C-terminal region of UBQLN2 is lysine-depleted, and we show that introducing lysine residues in this region leads to its E6AP-dependent degradation. The UBL domain critically stabilizes UBQLN2 and protects it from proteasomal degradation. Fusion of ubiquitin to the UBQLN2 N-terminus stabilizes UBQLN2 and increases its propensity for locating in puncta, indicating that ubiquitylation of the UBQLN2 UBL domain regulates abundance and localization.

biochemistry↗

AbDesign: Database of point mutants of antibodies with associated structures reveals poor generalization of binding predictions from machine learning models.

Antibodies are naturally evolved molecular recognition scaffolds that can bind a variety of surfaces. Their designability is crucial to the development of biologics, with computational methods holding promise in accelerating the delivery of medicines to the clinic. Modeling antibody-antigen recognition is prohibitively difficult, with data paucity being one of the biggest hurdles. Current affinity datasets comprise a small number of experimental measurements, which are often not standardized between molecules. Here, we address these issues by creating a dataset of seven antigens with two antibodies each, for which we introduce a heterogeneous set of mutations to the CDR-H3 measured by ELISA. Each of the parental complexes has a known crystal structure. We perform benchmarking of state-of-the-art affinity prediction algorithms to gauge their effectiveness. Current computational methods exhibit significant limitations in accurately predicting the effects of single-point mutations. In contrast, the older empirical, physics-based method FoldX, performs well in identifying mutants that retain binding. These findings highlight the need for more resources like the one presented here -- large, molecularly diverse, and experimentally consistent datasets.

bioinformatics↗