Search bioRxiv⌕ Search

Biology subjects

Jonsson, N.

Publications and source records attributed to Jonsson, N..

6 recordsLinked to original sources

Comprehensive degron mapping in human transcription factors

Gene expression is regulated by the targeted degradation of transcription factors through the ubiquitin-proteasome system. Transcription factors destined for degradation are recognized by E3 ubiquitin-protein ligases through short motifs termed degrons, embedded within the sequence. In this study, we systematically map degrons in all 1,626 human transcription factors. We find thousands of both known and previously unidentified degrons and characterize their sequence properties. Degrons placed within exposed and intrinsically disordered regions regulate the cellular abundance of the transcription factors, while the most common somatic mutations that are linked to skin cutaneous melanoma lead to unfolding and exposure of a buried degron in zinc fingers. We present examples of compartment specific degrons and demonstrate that variant effects in transcription factors correlate with degron potency. Finally, we show that while >60% of all predicted transcriptional activation domains overlap with strong degrons, acidic residues within the remaining transactivating regions counter the degron potency.

biochemistry↗

A complete map of human cytosolic degrons and their relevance for disease

Degrons are short protein segments that target proteins for degradation via the ubiquitin-proteasome system and thus ensure timely removal of signaling proteins and clearance of misfolded proteins from the intracellular space. Here, we describe a systematic screen for degrons in the human cytosol. We determine degron potency of >200,000 different 30-residue tiles from more than 5,000 cytosolic human proteins with 99.7% coverage. In total, 19.1% of the tiles function as strong degrons, 30.4% as intermediate degrons, while 50.5% did not display degron properties. The vast majority of the degrons are dependent on the E1 ubiquitin-activating enzyme and the proteasome but independent of autophagy. The results reveal both known and novel degron motifs, both internal as well as at the C-terminus. Mapping the degrons onto protein structures, predicted by AlphaFold2, revealed that most of the degrons are located in buried regions, indicating that they only become active upon unfolding or misfolding. Training of a machine learning model allowed us to probe the degron properties further and predict the cellular abundance of missense variants that operate by forming degrons in exposed and disordered protein regions, thus providing a mechanism of pathogenicity for germline coding variants at such positions.

biochemistry↗

Decoding molecular mechanisms for loss of function variants in the human proteome

Proteins play a critical role in cellular function by interacting with other biomolecules; missense variants that cause loss of protein function can lead to a broad spectrum of genetic disorders. While much progress has been made on predicting which missense variants may cause disease, our ability to predict the underlying molecular mechanisms remain limited. One common mechanism is that missense variants cause protein destabilization resulting in decreased protein abundance and loss of function, while other variants directly disrupt key interactions with other molecules. We have here leveraged machine-learning models for protein sequence and structure to disentangle effects on protein function and abundance, and applied our resulting model to all missense variants in the human proteome. We find that approximately half of all missense variants that lead to loss of function and disease do so because they disrupt protein stability. We predicted functionally important positions in all human proteins and found that they cluster on protein structures and are often found on the protein surface. Our work provides a resource for interpreting both predicted and experimental variant effects across the human proteome, and a mechanistic starting point for developing therapies towards genetic diseases.

bioinformatics↗

A joint embedding of protein sequence and structure enables robust variant effect predictions

The ability to predict how amino acid changes may affect protein function has a wide range of applications including in disease variant classification and protein engineering. Many existing methods focus on learning from patterns found in either protein sequences or protein structures. Here, we present a method for integrating information from protein sequences and structures in a single model that we term SSEmb (Sequence Structure Embedding). SSEmb combines a graph representation for the protein structure with a transformer model for processing multiple sequence alignments, and we show that by integrating both types of information we obtain a variant effect prediction model that is more robust to cases where sequence information is scarce. Furthermore, we find that SSEmb learns embeddings of the sequence and structural properties that are useful for other downstream tasks. We exemplify this by training a downstream model to predict protein-protein binding sites at high accuracy using only the SSEmb embeddings as input. We envisage that SSEmb may be useful both for zero-shot predictions of variant effects and as a representation for predicting protein properties that depend on protein sequence and structure.

bioinformatics↗

Conformational ensembles of the human intrinsically disordered proteome: Bridging chain compaction with function and sequence conservation

Intrinsically disordered proteins and regions (collectively IDRs) are pervasive across proteomes in all kingdoms of life, help shape biological functions, and are involved in numerous diseases. IDRs populate a diverse set of transiently formed structures, yet defy commonly held sequence-structure-function relationships. Recent developments in protein structure prediction have led to the ability to predict the three-dimensional structures of folded proteins at the proteome scale, and have enabled large-scale studies of structure-function relationships. In contrast, knowledge of the conformational properties of IDRs is scarce, in part because the sequences of disordered proteins are poorly conserved and because only few have been characterized experimentally. We have developed an efficient model to generate conformational ensembles of IDRs, and thereby to predict their conformational properties from sequence only. Here, we applied this model to simulate all IDRs of the human proteome. Examining conformational ensembles of 29,998 IDRs, we show how chain compaction is correlated with cellular function and localization, including in different types of biomolecular condensates. We train a model to predict compaction from sequence and use this to show conservation of structural properties across orthologs. Our results recapitulate observations from previous studies of individual protein systems, and enable us to study the relationship between sequence, conservation, conformational ensembles, biological function and disease variants at the proteome scale.

biophysics↗

Rapid protein stability prediction using deep learning representations

Predicting the thermodynamic stability of proteins is a common and widely used step in protein engineering, and when elucidating the molecular mechanisms behind evolution and disease. Here, we present RaSP, a method for making rapid and accurate predictions of changes in protein stability by leveraging deep learning representations. RaSP performs on-par with biophysics-based methods and enables saturation mutagenesis stability predictions in less than a second per residue. We use RaSP to calculate [~] 300 million stability changes for nearly all single amino acid changes in the human proteome, and examine variants observed in the human population. We find that variants that are common in the population are substantially depleted for severe destabilization, and that there are substantial differences between benign and pathogenic variants, highlighting the role of protein stability in genetic diseases. RaSP is freely available--including via a Web interface--and enables large-scale analyses of stability in experimental and predicted protein structures.

biophysics↗