Search bioRxiv⌕ Search

Biology subjects

Viknander, S.

Publications and source records attributed to Viknander, S..

4 recordsLinked to original sources

The Role of Metabolism in Shaping Enzyme Structures Over 400 Million Years of Evolution

The functions of cells and proteins depend on their biochemical microenvironment. To understand how biochemical constraints shaped protein structural evolution, we coupled the extensive genetic and metabolic data from the Saccharomycotina subphylum with the capability of AlphaFold2 to systematically predict protein structures from sequence. Determining how 11,269 enzyme structures catalysing 361 different metabolic reactions evolved over 400 million years alongside their molecular functions, we report that metabolism has shaped the structural evolution of enzymes at different levels: the organisms overall metabolism; the topological organisation of the metabolic network; and each enzymes molecular properties. For example, structural evolution depends on each enzymes reaction mechanism, on the variability rather than the amount of metabolic flux, and on biosynthetic cost. Evolutionary cost-optimization is stronger on highly abundant enzymes and acts differently on different structural domains, with the exception of small-molecule binding sites, which are prioritised over other structural domains and lack cost-optimisation. Finally, while enzyme surfaces are less constrained, surface residues can also be exposed to positive selection for the co-evolution of protein-protein interaction sites. Accessing AlphaFolds power to predict protein structures systematically and across species barriers, facilitating the integration of protein structures with functional genomics, we were thus able to map biological constraints which shape protein structural evolution at scale and over long timelines.

systems biology↗

The amino acid sequence determines protein abundance through its conformational stability and reduced synthesis cost.

Understanding what drives protein abundance is essential to biology, medicine, and biotechnology. Driven by evolutionary selection, the amino acid sequence is tailored to meet the required abundance of proteomes, underscoring the intricate relationship between sequence and functional demand. Yet, the specific role of amino acid sequences in determining proteome abundance remains elusive. Here, we demonstrate that the amino acid sequence predicts abundance by shaping a proteins conformational stability. We show that increasing the abundance provides metabolic cost benefits, underscoring the evolutionary advantage of maintaining a highly abundant and stable proteome. Specifically, using a deep learning model (BERT), we predict 56% of protein abundance variation in Saccharomyces cerevisiae solely based on amino acid sequence. The model reveals latent factors linking sequence features to protein stability. To probe these relationships, we introduce MGEM (Mutation Guided by an Embedded Manifold), a methodology for guiding protein abundance through sequence modifications. We find that mutations increasing abundance significantly alter protein polarity and hydrophobicity, underscoring a connection between protein stability and abundance. Through molecular dynamics simulations and in vivo experiments in yeast, we confirm that abundance-enhancing mutations result in longer-lasting and more stable protein expression. Importantly, these sequence changes also reduce metabolic costs of protein synthesis, elucidating the evolutionary advantage of cost-effective, high-abundance, stable proteomes. Our findings support the role of amino acid sequence as a pivotal determinant of protein abundance and stability, revealing an evolutionary optimization for metabolic efficiency.

evolutionary biology↗

Computational Scoring and Experimental Evaluation of Enzymes Generated by Neural Networks

In recent years, generative protein sequence models have been developed to sample novel sequences. However, predicting whether generated proteins will fold and function remains challenging. We evaluate computational metrics to assess the quality of enzyme sequences produced by three contrasting generative models: ancestral sequence reconstruction, a generative adversarial network, and a protein language model. Focusing on two enzyme families, we expressed and purified over 440 natural and generated sequences with 70-90% identity to the most similar natural sequences to benchmark computational metrics for predicting in vitro enzyme activity. Over three rounds of experiments, we developed a computational filter that improved experimental success rates by 44-100%. Surprisingly, neither sequence identity to natural sequences nor AlphaFold2 residue-confidence scores were predictive of enzyme activity. The proposed metrics and models will drive protein engineering research by serving as a benchmark for generative protein sequence models and helping to select active variants to test experimentally.

biochemistry↗

Learning deep representations of enzyme thermal adaptation

Temperature is a fundamental environmental factor that shapes the evolution of organisms. Learning thermal determinants of protein sequences in evolution thus has profound significance for basic biology, drug discovery, and protein engineering. Here, we use a dataset of over 3 million enzymes labeled with optimal growth temperatures (OGT) of their source organisms to train a deep neural network model (DeepET). The protein-temperature representations learned by DeepET provide a temperature-related statistical summary of protein sequences and capture structural properties that affect thermal stability. For prediction of enzyme optimal catalytic temperatures and protein melting temperatures via a transfer learning approach, our DeepET model outperforms classical regression models trained on rationally designed features and other recent deep-learning-based representations. DeepET thus holds promise for understanding enzyme thermal adaptation and guiding the engineering of thermostable enzymes.

bioinformatics↗