Search bioRxiv⌕ Search

Biology subjects

Ertelt, M.

Publications and source records attributed to Ertelt, M..

2 recordsLinked to original sources

HyperMPNN - A general strategy to design thermostable proteins learned from hyperthermophiles

Stability is a key factor to enable the use of recombinant proteins in therapeutic or biotechnological applications. Deep learning protein design approaches like ProteinMPNN have shown strong performance both in creating novel proteins or stabilizing existing ones. However, it is unlikely that the stability of the designs will significantly exceed that of the natural proteins in the training set, which are biophysically only marginally stable. Therefore, we collected predicted protein structures from hyperthermophiles, which differ substantially in their amino acid composition from mesophiles. Notably, ProteinMPNN fails to recover their unique amino acid composition. Here we show that a retrained network on predicted proteins from hyperthermophiles, termed HyperMPNN, not only recovers this unique amino acid composition but can also be applied to proteins from non-hyperthermophiles. Using this novel approach on a protein nanoparticle with a melting temperature of 65{degrees}C resulted in designs remaining stable at 95{degrees}C. In conclusion, we created a new way to design highly thermostable proteins through self-supervised learning on data from hyperthermophiles.

bioinformatics↗

Self-supervised machine learning methods for protein design improve sampling, but not the identification of high-fitness variants

Machine learning (ML) is changing the world of computational protein design, with data- driven methods surpassing biophysical-based methods in experimental success rates. However, they are most often reported as case studies, lack integration and standardization across platforms, and are therefore hard to objectively compare. In this study, we established a streamlined and diverse toolbox for methods that predict amino acid probabilities inside the Rosetta software framework that allows for the side-by-side comparison of these models. Subsequently, existing protein fitness landscapes were used to benchmark novel self- supervised machine learning methods in realistic protein design settings. We focused on the traditional problems of protein sequence design: sampling and scoring. A major finding of our study is that novel ML approaches are better at purging the sampling space from deleterious mutations. Nevertheless, scoring resulting mutations without model fine-tuning showed no clear improvement over scoring with Rosetta. This study fills an important gap in the field and allows for the first time a comprehensive head-to-head comparison of different ML and biophysical methods. We conclude that ML currently acts as a complement to, rather than a replacement for, biophysical methods in protein design.

bioinformatics↗