Search bioRxiv⌕ Search

Biology subjects

Degn, K.

Publications and source records attributed to Degn, K..

4 recordsLinked to original sources

PDBminer to Find and Annotate Protein Structures for Computational Analysis

Structural bioinformatics and molecular modeling of proteins strongly depend on the protein structure selected for investigation. The choice of protein structure relies on direct application from the Protein Data Bank (PDB), homology- or de-novo modeling. Recent de-novo models, such as AlphaFold2, require little preprocessing and omit the need to navigate the many parameters of choosing an experimentally determined model. Yet, the experimentally determined structure still has much to offer, why it should be of interest to the community to ease the choice of experimentally determined models. We provide an open-source software package, PDBminer, to mine both the AlphaFold Database (AlphaFoldDB) and the PDB based on search criteria set by the user. This tool provides an up-to-date, quality-ranked table of structures applicable for further research. PDBminer provides an overview of the available protein structures to one or more input proteins, parallelizing the runs if multiple cores are specified. The output table reports the coverage of the protein structures aligned to the UniProt sequence, overcoming numbering differences in PDB structures, and providing information regarding model quality, protein complexes, ligands, and nucleotide binding. The PDBminer2coverage and PDBminer2network tools assist in visualizing the results. We suggest that PDBminer can be applied to overcome the tedious task of choosing a PDB structure without losing the wealth of additional information available in the PDB. As developers, we will guarantee the introduction of new functionalities, assistance, training of new contributors, and package maintenance. The package is available at http://github.com/ELELAB/PDBminer.

bioinformatics↗

TRAP1 S-nitrosylation as a model of population-shift mechanism to study the effects of nitric oxide on redox-sensitive oncoproteins.

S-nitrosylation is a post-translational modification in which nitric oxide (NO) binds to the thiol group of cysteine, generating an S-nitrosothiol (SNO) adduct. S-nitrosylation has different physiological roles, and its alteration has also been linked to a growing list of pathologies, including cancer. SNO can affect the function and stability of different proteins, such as the mitochondrial chaperone TRAP1. Interestingly, the SNO site (C501) of TRAP1 is in the proximity of another cysteine (C527). This feature suggests that the S-nitrosylated C501 could engage in a disulfide bridge with C527 in TRAP1, resembling the well-known ability of S-nitrosylated cysteines to resolve in disulfide bridge with vicinal cysteines. We used enhanced sampling simulations and in-vitro biochemical assays to address the structural mechanisms induced by TRAP1 S-nitrosylation. We showed that the SNO site induces conformational changes in the proximal cysteine and favors conformations suitable for disulfide-bridge formation. We explored 4172 known S-nitrosylated proteins using high-throughput structural analyses. Furthermore, we carried out coarse-grain simulations of 44 proteins to account for protein dynamics in the analyses. This resulted in the identification of up to 1248 examples of proximal cysteines which could sense the redox state of the SNO site, opening new perspectives on the biological effects of redox switches. In addition, we devised two bioinformatic workflows (https://github.com/ELELAB/SNO_investigation_pipelines) to identify proximal or vicinal cysteines for a SNO site with accompanying structural annotations. Finally, we analyzed mutations in tumor suppressor or oncogenes in connection with the conformational switch induced by S-nitrosylation. We classified the variants as neutral, stabilizing, or destabilizing with respect to the propensity to be S-nitrosylated and to undergo the population-shift mechanism. The methods applied here provide a comprehensive toolkit for future high-throughput studies of new protein candidates, variant classification, and a rich data source for the research community in the NO field.

biochemistry↗

MAVISp: Multi-layered Assessment of VarIants by Structure for proteins

The role of genomic variants in disease has expanded significantly with the advent of advanced sequencing techniques. The rapid increase in identified genomic variants has led to many variants being classified as Variants of Uncertain Significance or as having conflicting evidence, posing challenges for their interpretation and characterization. Additionally, current methods for predicting pathogenic variants often lack insights into the underlying molecular mechanisms. Here, we introduce MAVISp (Multi-layered Assessment of VarIants by Structure for proteins), a modular structural framework for variant effects, accompanied by a web server (https://services.healthtech.dtu.dk/services/MAVISp-1.0/) to enhance data accessibility, consultation, and reusability. MAVISp currently provides data over 1000 proteins, encompassing more than eight million variants. A team of biocurators regularly analyzes and updates protein entries using standardized workflows, incorporating free energy calculations or biomolecular simulations. We illustrate the utility of MAVISp through selected case studies. The framework facilitates the analysis of variant effects at the protein level and has the potential to advance the understanding and application of mutational data in disease research.

bioinformatics↗

RosettaDDGPrediction for high-throughput mutational scans: from stability to binding

Reliable prediction of free energy changes upon amino acidic substitutions ({Delta}{Delta}Gs) is crucial to investigate their impact on protein stability and protein-protein interaction. Moreover, advances in experimental mutational scans allow high-throughput studies thanks to sophisticated multiplex techniques. On the other hand, genomics initiatives provide a large amount of data on disease-related variants that can benefit from analyses with structure-based methods. Therefore, the computational field should keep the same pace and provide new tools for fast and accurate high-throughput calculations of {Delta}{Delta}Gs. In this context, the Rosetta modeling suite implements effective approaches to predict the change in the folding free energy in a protein monomer upon amino acid substitutions and calculate the changes in binding free energy in protein complexes. Their application can be challenging to users without extensive experience with Rosetta. Furthermore, Rosetta protocols for {Delta}{Delta}G prediction are designed considering one variant at a time, making the setup of high-throughput screenings cumbersome. For these reasons, we devised RosettaDDGPrediction, a customizable Python wrapper designed to run free energy calculations on a set of amino acid substitutions using Rosetta protocols with little intervention from the user. RosettaDDGPrediction assists with checking whether the runs are completed successfully aggregates raw data for multiple variants, and generates publication-ready graphics. We showed the potential of the tool in selected case studies, including variants of unknown significance found in children who developed cancer, proteins with known experimental unfolding {Delta}{Delta}Gs values, interactions between target proteins and a disordered functional motif, and phospho-mimetic variants. RosettaDDGPrediction is available, free of charge and under GNU General Public License v3.0, at https://github.com/ELELAB/RosettaDDGPrediction.

bioinformatics↗