Search bioRxiv⌕ Search

Biology subjects

Stamkopoulos, E.

Publications and source records attributed to Stamkopoulos, E..

3 recordsLinked to original sources

Generative design of antibody Fc-variants with synthetic and programmable functional profiles

Beyond antigen recognition, antibodies direct diverse immune effector functions through their constant (Fc) domain. While the Fc domain is central to antibody biology and therapeutic efficacy, our understanding of how Fc sequence encodes function remains limited, as most of Fc sequence space has not been experimentally mapped or linked to Fc-receptor engagement. Furthermore, the extensive overlap in Fc-receptor binding sites on the Fc domain has impeded efforts to engineer antibodies with tailored, multi-receptor engagement profiles that can precisely control downstream immunity. Here we introduce a novel framework for Fc engineering that integrates protein engineering with deep learning to rationally predict and engineer antibody Fc function. Using a yeast-based, aglycosylated Fc display system, we performed deep mutational scanning across the entire human IgG1 Fc domain, allowing the rational design of a diverse combinatorial library of more than 108 Fc-variants. This library was sorted based on binding to a panel of eight canonical Fc-receptors, and the resulting populations were deep sequenced to generate a high-quality dataset comprising millions of unique Fc sequences annotated with their respective Fc-receptor binding profiles. Deep learning-based classifiers trained on this dataset accurately predicted Fc-receptor binding activity from Fc sequence across all Fc-receptors tested. We further developed FcGPT, a domain-specific autoregressive protein language model pre-trained on over three million unique Fc sequences, and refined by post-training through reinforcement learning with experimental feedback (RLXF) and synthetic verifiers. FcGPT enables the computational design of novel Fc-variants with user-defined Fc-receptor binding profiles, providing a foundational tool for understanding and programming antibody-mediated immunity.

bioengineering↗

Dissecting serum polyclonal antibody escape to SARS-CoV-2 variants by deep mutational learning

The rapid emergence of SARS-CoV-2 variants harboring multiple receptor-binding domain (RBD) mutations continues to challenge the efficacy of vaccines and antibody therapeutics. While deep mutational scanning (DMS) has been instrumental in mapping single-mutation effects on antibody binding and immune escape, it remains limited in its ability to assess combinatorial mutational landscapes. Here, we extend deep mutational learning (DML), a method integrating combinatorial mutagenesis, yeast surface display, deep sequencing, and machine learning, to analyze serum polyclonal antibody escape. Human sera from COVID-19-vaccinated individuals were screened against diverse RBD variant libraries, validating 300 serum-variant interactions across 10 individuals by comparing predicted and observed binding to 30 RBD variants, confirming accurate mapping of binding and escape profiles. Model performance remained consistent across machine learning architectures, suggesting that serum binding and escape are governed by distinct, localized RBD sequence features. Notably, escape profiles were highly individualized, whereas binding signatures were more conserved, reflecting convergent epitope targeting. This approach highlights the potential of DML to generalize beyond observed RBD variants, assess cohort-specific immune breadth, and inform vaccine and therapeutic design in the face of viral evolution.

immunology↗

PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering

Protein language models (PLMs) have emerged as a useful resource for protein engineering applications. Transfer learning (TL) leverages pre-trained parameters to extract features to train machine learning models or adjust the weights of PLMs for novel tasks via fine-tuning through back-propagation. TL methods have shown potential for enhancing protein predictions performance when paired with PLMs, however there is a notable lack of comparative analyses that benchmark TL methods applied to state-of-the-art PLMs, identify optimal strategies for transferring knowledge and determine the most suitable approach for specific tasks. Here, we report PLMFit, a benchmarking study that combines, three state-of-the-art PLMs (ESM2, ProGen2, ProteinBert), with three TL methods (feature extraction, low-rank adaptation, bottleneck adapters) for five protein engineering datasets. We conducted over >3,150 in silico experiments, altering PLM sizes and layers, TL hyperparameters and different training procedures. Our experiments reveal three key findings: (i) utilizing a partial fraction of PLM for TL does not detrimentally impact performance, (ii) the choice between feature extraction and fine-tuning is primarily dictated by the amount and diversity of data and (iii) fine-tuning is most effective when generalization is necessary and only limited data is available. We provide PLMFit as an open-source software package, serving as a valuable resource for the scientific community to facilitate the feature extraction and fine-tuning of PLMs for various applications. ONE SENTENCE SUMMARYPLMFit is a comparative analysis aimed at identifying the most effective strategies for transfer knowledge from protein language models by benchmarking fine-tuning techniques on a range of protein engineering tasks.

bioinformatics↗