Search bioRxivSearch

Biology subjects

Lindorff-Larsen, K.

Publications and source records attributed to Lindorff-Larsen, K..

8 recordsLinked to original sources

Side chain to main chain hydrogen bonds stabilize polyglutamine helices in transcription factors

Polyglutamine (polyQ) tracts are regions of low sequence complexity of variable length found in more than one hundred human proteins. These tracts are frequent in activation domains of transcription factors and their length often correlates with transcriptional activity. In addition, in nine proteins, tract elongation beyond specific thresholds causes polyQ disorders. To study the structural basis of the association between tract length, transcriptional activity and disease, here we addressed how the conformation of the polyQ tract of the androgen receptor (AR), a transcription factor associated with the polyQ disease spinobulbar muscular atrophy (SBMA), depends on its length. We found that the tract folds into a helical structure stabilized by unconventional hydrogen bonds between glutamine side chains and main chain carbonyl groups. These bonds are bifurcate with the conventional main chain to main chain hydrogen bonds stabilizing -helices. In addition, since tract elongation provides additional interactions, the helicity of the polyQ tract directly correlates with its length. These findings suggest a plausible rationale for the association between polyQ tract length and AR transcriptional activity and have implications for establishing the mechanistic basis of SBMA.

biophysics

Monte Carlo Sampling of Protein Folding by Combining an All-Atom Physics-Based Model with a Native State Bias

Energy landscape theory suggests that native interactions are a major determinant of the folding mechanism of a protein. Thus, structure-based (G[o]) models have, aided by coarse-graining techniques, shown great success in capturing the mechanisms of protein folding and conformational changes. In certain cases, however, non-native interactions and atomic details are also essential to describe the protein dynamics, prompting the development of a variety of structure-based models which include non-native interactions, and differentiate between different types of attractive potentials. Here, we describe an all-protein-atom hybrid model, termed ProfasiGo, that integrates an implicit solvent all-atom physics-based model (called Profasi) and a structure-based G[o] potential, and its implementation in two software packages (PHAISTOS and ProFASi) that are developed for Monte Carlo sampling of protein molecules. We apply the ProfasiGo model to study the folding free energy landscapes of four topologically similar proteins, one of which can be folded by the simplified potential Profasi, and two that have been folded by explicit solvent, all-atom molecular dynamics simulations with the CHARMM22* force field. Our results reveal that the hybrid ProfasiGo model is able to capture many of the details present in the physics-based potentials, while retaining the advantages of G[o] models for sampling and guiding to the native state. We expect that the model will be widely applicable to study the folding of more complex proteins, or to study conformational dynamics and integration with experimental data.

biophysics

Analyze Nucleic Acids Structures and Trajectories with Barnaba.

RNA molecules are highly dynamic systems characterized by a complex interplay between sequence, structure, dynamics, and function. Molecular simulations can potentially provide powerful insights into the nature of these relationships. The analysis of structures and molecular trajectories of nucleic acids can be non-trivial because it requires processing very high-dimensional data that are not easy to visualize and interpret.\n\nHere we introduce Barnaba, a Python library aimed at facilitating the analysis of nucleic acids structures and molecular simulations. The software consists of a variety of analysis tools that allow the user to i) calculate distances between three-dimensional structures using different metrics, ii) back-calculate experimental data from three-dimensional structures, iii) perform cluster analysis and dimensionality reductions, iv) search three-dimensional motifs in PDB structures and trajectories and v) construct elastic network models (ENM) for nucleic acids and nucleic acids-protein complexes.\n\nIn addition, Barnaba makes it possible to calculate torsion angles, pucker conformations and to detect base-pairing/base-stacking interactions. Barnaba produces graphics that conveniently visualize both extended secondary structure and dynamics for a set of molecular conformations. The software is available as a command-line tool as well as a library, and supports a variety of 1le formats such as PDB, dcd and xtc 1les. Source code, documentation and examples are freely available at https://github.com/srnas/barnaba under GNU GPLv3 license.

bioinformatics

Enhancing coevolution-based contact prediction by imposing structural self-consistency of the contacts

Based on the development of new algorithms and growth of sequence databases, it has recently become possible to build robust and informative higher-order statistical sequence models based on large sets of aligned protein sequences. By disentangling direct and indirect effects, such models have proven useful to assess phenotypic landscapes, determine protein-protein interaction sites, and in de novo structure prediction. In the context of structure prediction, the sequence models are used to find pairs of residues that co-vary during evolution, and hence are likely to be in spatial proximity in the functional native protein. The accuracy of these algorithms, however, drop dramatically when the number of sequences in the alignment is small, and thus the highest ranking pairs may include a substantial number of false positive predictions. We have developed a method that we termed CE-YAPP (CoEvolution-YAPP), that is based on YAPP (Yet Another Peak Processor), which has been shown to solve a similar problem in NMR spectroscopy. By simultaneously performing structure prediction and contact assignment, CE-YAPP uses structural self-consistency as a filter to remove false positive contacts. At the same time CE-YAPP solves another problem, namely how many contacts to choose from the ordered list of covarying amino acid pairs. Our results show that CE-YAPP consistently and substantially improves contact prediction from multiple sequence alignments, in particular for proteins that are difficult targets. We further show that CE-YAPP can be integrated with many different contact prediction methods, and thus will benefit also from improvements in algorithms for sequence analyses. Finally, we show that the structures determined from CE-YAPP are also in better agreement with those determined using traditional methods in structural biology.\n\nAuthor summaryHomologous proteins generally have similar functions and three-dimensional structures. This in turn means that it is possible to extract structural information from a detailed analysis of a multiple sequence alignment of a protein sequence. In particular, it has been shown that global statistical analyses of such sequence alignments allows one to find pairs of residues that have covaried during evolution, and that such pairs are likely to be in close contact in the folded protein structure. Although these insights have led to important developments in our ability to predict protein structures, these methods generally result in many false positive contacts predicted when the number of homologous sequences is not large. To deal with this issue, we have developed CE-YAPP, a method that can take a noisy set of predicted contacts as input and robustly detect many incorrectly predicted contacts within these. More specifically, our method performs simultaneous structure prediction and contact assignment so as to use structural self-consistency as a filter for erroneous predictions. In this way, CE-YAPP improves contact and structure predictions, and thus advances our ability to extract structural information from analyses of the evolutionary record of a protein.

bioinformatics

Deep mutational scanning by FACS-sorting of encapsulated E. coli micro-colonies

We present a method for high-throughput screening of protein variants where the signal is enhanced by micro-encapsulation of single cells into 20-30 m agarose beads. Cells inside beads are propagated using standard agitation in liquid media and grow clonally into micro-colonies harboring several hundred bacteria. We have, as a proof-of-concept, analyzed random amino acid substitutions in the five C-terminal {beta}-strands of the Green Fluorescent Protein (GFP). Starting from libraries of variants, each bead represents a clonal line of cells that can be separated by Fluorescence Activated Cell Sorting (FACS). Pools representing collections of individual variants with desired properties are subsequently analyzed by deep sequencing. Notably, the encapsulation approach described holds the potential for high-throughput analysis of systems where the fluorescence signal from a single cell is insufficient for detection. Fusion to GFP, or use of fluorogenic substrates, allows coupling protein levels or activity to sequence for a wide range of proteins. Here we analyzed more than 10,000 individual variants to gauge the effect of mutations on GFP-fluorescence. In the mutated region, we observed virtually all amino acid substitutions that are accessible by single nucleotide exchange. Lastly, we assessed the performance of biophysical protein stability predictors, FoldX and Rosetta, in predicting the outcome of the experiment. Both tools display good performance on average, suggesting that loss of thermodynamic stability is a key mechanism for the observed variation of the mutants. This, in turn, suggests that deep mutational scanning datasets may be used to more efficiently fine-tune such predictors, especially for mutations poorly covered by current biophysical data.

bioengineering

Molecular dynamics ensemble refinement of the heterogeneous native state of NCBD using chemical shifts and NOEs

Many proteins display complex dynamical properties that are often intimately linked to their biological functions. As the native state of a protein is best described as an ensemble of confor-mations, it is important to be able to generate models of native state ensembles with high accuracy. Due to limitations in sampling efficiency and force field accuracy it is, however, challenging to obtain accurate ensembles of protein conformations by the use of molecular simulations alone. Here we show that dynamic ensemble refinement, which combines an accurate atomistic force field with commonly available nuclear magnetic resonance (NMR) chemical shifts and NOEs, can provide a detailed and accurate description of the conformational ensemble of the native state of a highly dynamic protein. As both NOEs and chemical shifts are averaged on timescales up to milliseconds, the resulting ensembles reflect the structural heterogeneity that goes beyond that probed e.g. by NMR relaxation order parameters. We selected the small protein domain NCBD as object of our study since this protein, which has been characterized experimentally in substantial detail, displays a rich and complex dynamical behaviour. In particular, the protein has been described as having a molten-globule like structure, but with a relatively rigid core. Our approach allowed us to describe the conformational dynamics of NCBD in solution, and to probe the structural heterogeneity resulting from both short- and long-time-scale dynamics by the calculation of order parameters on different time scales. These results illustrate the usefulness of our approach since they show that NCBD is rather rigid on the nanosecond timescale, but interconverts within a broader ensemble on longer timescales, thus enabling the derivation of a coherent set of conclusions from various NMR experiments on this protein, which could otherwise appear in contradiction with each other.

biophysics

How well do force fields capture the strength of salt bridges in proteins?

Salt bridges form between pairs of ionisable residues in close proximity and are important interactions in proteins. While salt bridges are known to be important both for protein stability, recognition and regulation, we still do not have fully accurate predictive models to assess the energetic contributions of salt bridges. Molecular dynamics simulations is one technique that may be used study the complex relationship between structure, solvation and energetics of salt bridges, but the accuracy of such simulations depend on the force field used. We have used NMR data on the B1 domain of protein G (GB1) to benchmark molecular dynamics simulations. Using enhanced sampling simulations, we calculated the free energy of forming a salt bridge for three possible ionic interactions in GB1. The NMR experiments showed that these interactions are either not formed, or only very weakly formed, in solution. In contrast, we show that the stability of the salt bridges is slightly overestimated in simulations of GB1 using six commonly used combinations of force fields and water models. We therefore conclude that further work is needed to refine our ability to model quantitatively the stability of salt bridges through simulations, and that comparisons between experiments and simulations will play a crucial role in furthering our understanding of this important interaction.

biophysics

Conformational Ensemble of RNA Oligonucleotides from Reweighted Molecular Simulations

We determine the conformational ensemble of four RNA tetranucleotides by using available nuclear magnetic spectroscopy data in conjunction with extensive atomistic molecular dynamics simulations. This combination is achieved by applying a reweighting scheme based on the maximum entropy principle. We provide a quantitative estimate for the population of different conformational states by considering different NMR parameters, including distances derived from nuclear Overhauser effect intensities and scalar coupling constants. We show the usefulness of the method as a general tool for studying the conformational dynamics of flexible biomolecules as well as for detecting inaccuracies in molecular dynamics force fields.

biophysics