Search bioRxiv⌕ Search

Biology subjects

Roda, S.

Publications and source records attributed to Roda, S..

2 recordsLinked to original sources

Efficient Design of Affilin(R) Protein Binders for HER3

Engineered scaffold-based proteins that bind to concrete targets with high affinity offer significant advantages over traditional antibodies in theranostic applications. Their development often relies on display methods, where large libraries of variants are physically contacted with the desired target protein and pools of binding variants can be selected. Herein, we use a combined artificial intelligence/physics-based computational framework and phage display approach to obtain ubiquitin based Affilin(R) proteins targeting the HER3 extracellular domain, a relevant tumor target. We demonstrate that the developed in silico pipeline can generate de novo Affilin(R) proteins with high experimental success rate using a small training set of sequences (<1000 sequences). The classical phage display yielded primary candidates with low nanomolar affinities. These binders could be further optimized by phage display and computational maturation alike. These combined efforts resulted in four HER3 ligands with high affinity, cell binding, and serum stability that have theranostic potential.

biochemistry↗

Efficient and accurate sequence generation with small-scale protein language models

Large Language Models (LLMs) have demonstrated exceptional capabilities in understanding contextual relationships, outperforming traditional methodologies in downstream tasks such as text generation and sentence classification. This success has been mirrored in the realm of protein language models (pLMs), where proteins are encoded as text via their amino acid sequences. However, the training of pLMs, which involves tens to hundreds of millions of sequences and hundreds of millions to billions of parameters, poses a significant computational challenge. In this study, we introduce a Small-Scale Protein Language Model (SS-pLM), a more accessible approach that requires training on merely millions of representative sequences, reducing the number of trainable parameters to 14.8M. This model significantly reduces the computational load, thereby democratizing the use of foundational models in protein studies. We demonstrate that the performance of our model, when fine-tuned to a specific set of sequences for generation, is comparable to that of larger, more computationally demanding pLM.

bioinformatics↗