Search bioRxiv⌕ Search

Biology subjects

Jung, M. D.

Publications and source records attributed to Jung, M. D..

3 recordsLinked to original sources

Accurate protein stability prediction for small domains using mega-scale experiments

Predicting absolute protein folding stability is a long-standing challenge in biophysics, with broad applications in protein design and in understanding genetic variation and evolution. Physics-based simulations have shown limited success at predicting stability and are often computationally intractable, and machine learning methods have been constrained by the lack of sufficiently large experimental datasets. We recently introduced cDNA display proteolysis, a cell-free approach that can measure folding stability for nearly one million protein domains in parallel. Here, we applied this method to measure stability for 1.8 million diverse protein domains 60-80 amino acids in length primarily taken from the MGnify metagenomic database and spanning over 200,000 sequence families. Using this new "MGnify Stability dataset", we developed the predictive models SaProt{Delta}G and ESM3{Delta}G, which accurately predict absolute folding stability for small domains with root mean squared error of 0.8 kcal/mol over a 6 kcal/mol range (Spearman rank correlation of 0.88). These predictors show high accuracy at predicting effects of substitutions, insertions, and deletions, successfully identify global trends toward higher stability in thermophilic organisms, and improve discrimination of stable and unstable computationally designed proteins. Our results illustrate how megascale biophysical measurements can complement existing evolutionary and structural data to enable accurate absolute stability prediction for small domains.

biophysics↗

Global Analysis of Aggregation Determinants in Small Protein Domains

Protein aggregation is an obstacle for engineering effective recombinant proteins for biotechnology and therapeutic applications. Predicting protein aggregation propensity remains challenging due to the complex interplay of sequence, structure, environmental factors, and external stress conditions, particularly for globular proteins. To understand the determinants of aggregation and improve its prediction, we quantified insoluble aggregation following high temperature and acidic stress in custom libraries of small protein domains (40-72 amino acids) using a high-throughput, in vitro, mass spectrometry-based method. In total, we quantified aggregation for 18,987 small protein domains, revealing diverse stress-dependent aggregation phenotypes that were consistent in different library contexts. We also found that aggregation measurements on individually purified proteins strongly correlated with high-throughput mixed-pool data (Pearsons r = 0.65-0.79), supporting the use of multiplexed approaches to study aggregation. Using machine learning, we identified sequence and structural features that correlate with aggregation and fine-tuned the protein language model SaProt, which explained 43-55% of the observed variation in a held-out test set of unrelated protein domains. Our model shows promising utility for engineering aggregation-resistant proteins, and our dataset serves as an important resource for developing improved models of protein aggregation.

biophysics↗

Large-scale discovery, analysis, and design of protein energy landscapes

All folded proteins continuously fluctuate between their low-energy native structures and higher energy conformations that can be partially or fully unfolded. These rare states influence protein function, interactions, aggregation, and immunogenicity, yet they remain far less understood than protein native states. Although native protein structures are now often predictable with impressive accuracy, conformational fluctuations and their energies remain largely invisible and unpredictable, and experimental challenges have prevented large-scale measurements that could improve machine learning and physics-based modeling. Here, we introduce a multiplexed experimental approach to analyze the energies of conformational fluctuations for hundreds of protein domains in parallel using intact protein hydrogen-deuterium exchange mass spectrometry. We analyzed 5,778 domains 28-64 amino acids in length, revealing hidden variation in conformational fluctuations even between sequences sharing the same fold and global folding stability. Site-resolved hydrogen exchange NMR analysis of 13 domains showed that these fluctuations often involve entire secondary structural elements with lower stability than the overall fold. Computational modeling of our domains identified structural features that correlated with the experimentally observed fluctuations, enabling us to design mutations that stabilized low-stability structural segments. Our dataset enables new machine learning-based analysis of protein energy landscapes, and our experimental approach promises to reveal these landscapes at unprecedented scale.

biophysics↗