Search bioRxiv⌕ Search

Biology subjects

Blaabjerg, L. M.

Publications and source records attributed to Blaabjerg, L. M..

3 recordsLinked to original sources

Supervised learning of protein variant effects across large-scale mutagenesis datasets

The increasing availability of data from multiplexed assays of variant effects (MAVEs) enables supervised model training against large quantities of experimental data to learn sequence-function relationships. Variant effect scores from MAVEs can, however, be influenced by the experimental method to create experiment-to-experiment differences in the mapping from molecular level-variant effects to MAVE readout, which presents a challenge for supervised learning across datasets. We here propose a framework for performing supervised learning with MAVE data that takes the influence of the experimental protocol into account, thus enabling variant effects to be learned across datasets produced in independent experiments. We apply the framework to train a model against variant effect scores collected with VAMP-seq, a MAVE technique that quantifies the steady-state cellular abundance of protein variants. We show that mapping variant abundance to VAMP-seq readout in a dataset-specific manner during model training improves the learned abundance model and moreover allows the learned model to predict variant effects on an interpretable scale. Our work highlights the importance of validating MAVE results with low-throughput methods to facilitate MAVE score interpretation and supervised model training.

biophysics↗

A joint embedding of protein sequence and structure enables robust variant effect predictions

The ability to predict how amino acid changes may affect protein function has a wide range of applications including in disease variant classification and protein engineering. Many existing methods focus on learning from patterns found in either protein sequences or protein structures. Here, we present a method for integrating information from protein sequences and structures in a single model that we term SSEmb (Sequence Structure Embedding). SSEmb combines a graph representation for the protein structure with a transformer model for processing multiple sequence alignments, and we show that by integrating both types of information we obtain a variant effect prediction model that is more robust to cases where sequence information is scarce. Furthermore, we find that SSEmb learns embeddings of the sequence and structural properties that are useful for other downstream tasks. We exemplify this by training a downstream model to predict protein-protein binding sites at high accuracy using only the SSEmb embeddings as input. We envisage that SSEmb may be useful both for zero-shot predictions of variant effects and as a representation for predicting protein properties that depend on protein sequence and structure.

bioinformatics↗

Rapid protein stability prediction using deep learning representations

Predicting the thermodynamic stability of proteins is a common and widely used step in protein engineering, and when elucidating the molecular mechanisms behind evolution and disease. Here, we present RaSP, a method for making rapid and accurate predictions of changes in protein stability by leveraging deep learning representations. RaSP performs on-par with biophysics-based methods and enables saturation mutagenesis stability predictions in less than a second per residue. We use RaSP to calculate [~] 300 million stability changes for nearly all single amino acid changes in the human proteome, and examine variants observed in the human population. We find that variants that are common in the population are substantially depleted for severe destabilization, and that there are substantial differences between benign and pathogenic variants, highlighting the role of protein stability in genetic diseases. RaSP is freely available--including via a Web interface--and enables large-scale analyses of stability in experimental and predicted protein structures.

biophysics↗