Search bioRxiv⌕ Search

Biology subjects

Dieckhaus, H.

Publications and source records attributed to Dieckhaus, H..

3 recordsLinked to original sources

Assessing substrate scope of the cyclodehydratase LynD by mRNA display-enabled machine learning models

Many of the biosynthetic pathways for ribosomal synthesized and post-translationally modified peptide (RiPP) natural products make use of multi-domain enzymes with separate recruitment and catalysis domains that separately bind and modify peptide substrates. This "division of labor" allows RiPP enzymes to use relatively open and promiscuous active sites to perform chemistry at multiple residues within a peptide substrate seemingly regardless of the surrounding context. Defining, measuring, and predicting the seemingly broad substrate promiscuity of RiPPs necessitates high throughput assays, capable of assessing activity against very large libraries of peptides. Using mRNA display, a high throughput peptide display technology, we examine the substrate promiscuity of the RiPP cyclodehydratase, LynD. The vast substrate profiling that can be done with mRNA display enables the construction of deep learning models for accurate prediction of substrate processing by LynD. These models further inform on epistatic interactions involved in enzymatic processing. This work will facilitate the further elucidation of other RiPP enzymes and enable their use in the modification of mRNA display libraries for selection of modified peptide-based inhibitors and therapeutics.

biochemistry↗

Protein stability models fail to capture epistatic interactions of double point mutations

There is strong interest in accurate methods for predicting changes in protein stability resulting from amino acid mutations to the protein sequence. Recombinant proteins must often be stabilized to be used as therapeutics or reagents, and destabilizing mutations are implicated in a variety of diseases. Due to increased data availability and improved modeling techniques, recent studies have shown advancements in predicting changes in protein stability when a single point mutation is made. Less focus has been directed toward predicting changes in protein stability when there are two or more mutations, despite the significance of mutation clusters for disease pathways and protein design studies. Here, we analyze the largest available dataset of double point mutation stability and benchmark several widely used protein stability models on this and other datasets. We identify a blind spot in how predictors are typically evaluated on multiple mutations, finding that, contrary to assumptions in the field, current stability models are unable to consistently capture epistatic interactions between double mutations. We observe one notable deviation from this trend, which is that epistasis-aware models provide marginally better predictions on stabilizing double point mutations. We develop an extension of the ThermoMPNN framework for double mutant modeling as well as a novel data augmentation scheme which mitigates some of the limitations in available datasets. Collectively, our findings indicate that current protein stability models fail to capture the nuanced epistatic interactions between concurrent mutations due to several factors, including training dataset limitations and insufficient model sensitivity. SignificanceProtein stability is governed in part by epistatic interactions between energetically coupled residues. Prediction of these couplings represents the next frontier in protein stability modeling. In this work, we benchmark protein stability models on a large dataset of double point mutations and identify previously overlooked limitations in model design and evaluation. We also introduce several new strategies to improve modeling of epistatic couplings between protein point mutations.

bioinformatics↗

Transfer learning to leverage larger datasets for improved prediction of protein stability changes

Amino acid mutations that lower a proteins thermodynamic stability are implicated in numerous diseases, and engineered proteins with enhanced stability are important in research and medicine. Computational methods for predicting how mutations perturb protein stability are therefore of great interest. Despite recent advancements in protein design using deep learning, in silico prediction of stability changes has remained challenging, in part due to a lack of large, high-quality training datasets for model development. Here we introduce ThermoMPNN, a deep neural network trained to predict stability changes for protein point mutations given an initial structure. In doing so, we demonstrate the utility of a newly released mega-scale stability dataset for training a robust stability model. We also employ transfer learning to leverage a second, larger dataset by using learned features extracted from a deep neural network trained to predict a proteins amino acid sequence given its three-dimensional structure. We show that our method achieves competitive performance on established benchmark datasets using a lightweight model architecture that allows for rapid, scalable predictions. Finally, we make ThermoMPNN readily available as a tool for stability prediction and design.

bioinformatics↗