Search bioRxiv⌕ Search

Biology subjects

Streit, J. O.

Publications and source records attributed to Streit, J. O..

9 recordsLinked to original sources

Transient tertiary structure in intrinsically disordered proteins revealed by multithermal enhanced sampling

Intrinsically disordered proteins populate heterogeneous conformational ensembles that are challenging to characterise. While all-atom molecular dynamics simulations can provide highly detailed insights into dynamic ensembles, achieving sufficient sampling remains difficult. Here, we show that On-the-fly Probability Enhanced Sampling (OPES) in the multithermal ensemble enables efficient generation of atomistic ensembles for disordered peptides and proteins ranging from 15 to 71 residues in length. Using the potential energy as a collective variable, OPES achieves multithermal sampling within a single simulation replica, without replica exchange or extensive parameter tuning. Across multiple systems, OPES yields reweighted ensembles broadly consistent with replica-exchange with solute tempering (REST2) and unbiased simulations, while accelerating convergence and enabling broader exploration of low-population conformational states. Applied to the intrinsically disordered transcriptional coactivator ACTR, OPES reveals transiently structured states in which multiple -helices involved in partner binding fold cooperatively and form tertiary contacts. These rare, partially structured conformations are reversibly sampled during the simulations, consistent with extensive NMR and SAXS data, and could facilitate folding-upon-binding through conformational selection. They may also represent viable targets for drug design or for engineering disordered proteins with customised conformational landscapes. More broadly, our results establish OPES multithermal sampling as a robust and accessible approach for uncovering rare, functionally relevant conformations in intrinsically disordered proteins.

biophysics↗

Advancing Protein Ensemble Predictions Across the Order-Disorder Continuum

While deep learning has transformed structure prediction for ordered proteins, intrinsically disordered proteins remain poorly predicted due to systematic underrepresentation in training data, despite constituting approximately 30% of eukaryotic proteomes. We introduce PeptoneBench, the first benchmark to enable systematic assessment of ensemble generators for both ordered and disordered proteins, integrating diverse experimental observables. Our analysis reveals that existing evaluation metrics exhibit systematic bias toward the structured spectrum of the proteome. Assessment of popular predictors (AlphaFold2, ESMFlow, Boltz2) confirms high accuracy on ordered proteins but shows performance degradation with increasing disorder. We further present PepTron, a flow-matching ensemble generator trained on data augmented with synthetic disordered protein ensembles. On our benchmark PepTron matches BioEmu on disordered regions while maintaining competitive accuracy on ordered protein benchmarks. Our data augmentation approach demonstrates that targeted training strategies can approach the performance of computationally expensive simulation-based methods, establishing a generalizable framework applicable to other protein generative models. All datasets, models, and code are openly available.

biophysics↗

Protein Dynamics at Different Timescales Unlock Access to Hidden Post-Translational Modification Sites

Post-translational modifications (PTMs) alter the proteome in response to intra- and extracellular signals, providing fundamental information processing in development, homeostasis and disease1-5. Here, we used proteome-wide structure predictions to investigate the structural context of protein modification sites and found nearly a fifth to be buried within proteins6. These cryptic sites within folded domains showed evolutionary conservation across species, as well as significantly less variability within the human population. We also found significantly more modification sites to be affected by recurrent cancer mutations and disease variants, with some variants appearing to mimic the modified or unmodified state, supporting their functionality and disease relevance. To understand how these buried sites become modified, we studied 23 phosphorylation sites across seven proteins using all-atom molecular dynamics simulations. The simulations revealed a range of mechanisms spanning broad timescales, ranging from picosecond fluctuations to seconds for proline isomerisation-gated phosphorylation sites. Notably, some sites remained fully buried in the native structure. Some conformational states with exposed cryptic sites resembled structures of co-translational protein folding intermediates formed on the ribosome7, indicating an interplay between protein biogenesis, folding, modifications, and signalling. We propose more broadly that protein conformational fluctuations may serve as structure-specific regulatory mechanisms by kinetically burying or exposing the unmodified, and in some cases even modified, sidechains, thereby expanding the functional versatility of proteins and enabling signal transduction across diverse timescales.

bioinformatics↗

Visualisation of translating ribosomes reveals the earliest steps of protein misfolding in human disease

The majority of cellular proteins must adopt a particular three-dimensional structure for function1. However, protein folding is a perilous journey due to competing polypeptide misfolding events which result in inactive structures. In this study, we examine the earliest steps of protein misfolding during the biosynthesis of alpha-1-antitrypsin, a secreted plasma protein whose misfolding results in organ disease. Using human cells, we find that, like co-translational protein folding, misfolding, assembly and biosynthesis are interconnected processes. At the molecular level misfolding of alpha-1-antitrypsin is initiated by a molten globule-like folding intermediate formed cotranslationally on the ribosome. The ribosomal complexes subsequently form assemblies by recruiting released proteins, inducing translational arrest. Our data also reveal that a pharmacological chaperone modulates this process. The existence of co- and post-translational (mis)folding and assembly pathways reveals how some proteins form functional complexes, has implications for the pathogenesis of conformational diseases, and suggests novel therapeutic avenues.

biochemistry↗

Structures of protein folding intermediates on the ribosome

The ribosome biases the conformations sampled by nascent polypeptide chains along folding pathways towards biologically active states. A hallmark of the co-translational folding (coTF) of many proteins are highly stable folding intermediates that are absent or only transiently populated off the ribosome, yet persist during translation well-beyond complete emergence of the domain from the ribosome exit tunnel. Intermediates are important for folding fidelity; however, their structures have remained elusive. Here, we have structurally characterised two coTF intermediates of an immunoglobulin-like domain by developing comprehensive 19F NMR analyses using chemical shifts, paramagnetic relaxation enhancement (PRE), and protein engineering. We integrated these experimental data with extensive molecular dynamics (MD) simulations to obtain atomistic structures of the folding intermediates on the ribosome. The resulting intermediate structures are distinguished by native-like folds initiated from either their N-or C-termini, and reveal parallel folding pathways, which are structurally conserved within the protein domain family, in contrast to their in vitro refolding mechanisms. By redirecting proteins to fold along hierarchical, parallel routes, the ribosome may promote efficient folding by avoiding kinetic traps, and regulate nascent chain assembly and targeting by auxiliary factors to maintain cellular proteostasis.

biophysics↗

The ribosome directs nascent chains through two folding-dependent pathways

During their vectorial biosynthesis on the ribosome, elongating nascent polypeptide chains explore a range of conformational states towards their biologically functional structure. However, this high structural heterogeneity has limited their observation at high-resolution. Here, we have used an integrated structural biology approach to explore the structures of the multi-domain immunoglobulin-like FLN5-6 during its biosynthesis, capturing early folding through to native folding. We developed an in-silico purification approach for cryo-EM of ribosome-nascent chain complexes (RNCs), and integrated the resulting cryo-EM maps with NMR spectroscopy and atomistic molecular dynamics (MD) simulations to produce experimentally reweighted structural ensembles of RNC. The resulting atomistic structures reveal insights into the orientational heterogeneity of the nascent chain and its dynamic interactions with the ribosome. In particular, we find that two distinct pathways exist for nascent polypeptides in the exit tunnel vestibule, influenced by their stage of biosynthesis, folding conformational state and ribosomal RNA helices lining the tunnel. Our systematic analysis of the structures of nascent proteins translation-stalled at multiple time-points provides insights into how the ribosome dynamically modulates its pathway out of the exit tunnel to regulate its folding and accessibility for auxiliary factors of other co-translational events.

biophysics↗

Amyloid forming human lysozyme intermediates are stabilised by non-native amide-π interactions

Mutational variants of human lysozyme cause a rare but fatal hereditary form of systemic amyloidosis by populating an intermediate state that self-assembles into amyloid fibrils. Despite its significance in lysozyme amyloidosis, the intermediate state has been recalcitrant to detailed structural investigation as it is only transiently and sparsely populated. Here, we investigated the intermediate state of a mutational variant of human lysozyme (I59T) using CEST and CPMG RD NMR at low pH. 15N CEST profiles probed the thermal unfolding of the native state into the denatured ensemble and revealed an additional state distinct from the two major states. Global fitting of 15N CEST and CPMG data provided kinetic and thermodynamic parameters for the exchange between all three states, characterising the intermediate state populated at 0.6%. 1H CEST data also confirmed the presence of the intermediate state displaying unusually high or low 1HN chemical shifts. To further investigate the structural details of the intermediate state we used molecular dynamics (MD) simulations, which recapitulated the experimentally observed folding pathway and free energy landscape. A high-energy intermediate state with a locally disordered {beta}-domain and C-helix was observed, revealing non-native hydrogen bonding and amide-{pi} interactions. These interactions account for the anomalous 1H chemical shifts and likely stabilise the transient intermediate state structure. Together, our NMR and MD data provide the first direct structural information on the intermediate state, offering insights into targeting lysozyme amyloidosis.

biophysics↗

Rational design of 19F NMR labelling sites to probe protein structure and interactions

Proteins are investigated in increasingly more complex biological systems, where 19F NMR is proving highly advantageous due to its high gyromagnetic ratio and background-free spectra. Its application has, however, been hindered by limited chemical shift dispersions and an incomprehensive relationship between chemical shifts and protein structure. We exploit the sensitivity of 19F chemical shifts to ring currents by designing labels with direct contact to a native or engineered aromatic ring. Fifty protein variants predicted by AlphaFold and molecular dynamics simulations show 80-90% success rates and direct correlations of their experimental chemical shifts with the magnitude of the engineered ring current. Our method consequently improves the chemical shift dispersion and through simple 1D experiments enables structural analyses of alternative conformational states, including ribosome-bound folding intermediates, and in-cell measurements of thermodynamics and protein-protein interactions. Our strategy thus provides a simple and sensitive tool to extract residue contact restraints from chemical shifts for previously intractable systems.

biophysics↗

CryoENsemble - a Bayesian approach for reweighting biomolecular structural ensembles using heterogeneous cryo-EM maps

Cryogenic electron microscopy (cryo-EM) has emerged as a central tool for the determination of structures of complex biological molecules. Accurately characterising the dynamics of such systems, however, remains a challenge. To address this, we introduce cryoENsemble, a method that applies Bayesian reweighing to conformational ensembles derived from molecular dynamics simulations to improve their agreement with cryo-EM data and extract dynamics information. We illustrate the use of cryoENsemble to determine the dynamics of the ribosome-bound state of the co-translational chaperone trigger factor (TF). We also show that cryoENsemble can assist with the interpretation of low-resolution, noisy or unaccounted regions of cryo-EM maps. Notably, we are able to link an unaccounted part of the cryo-EM map to the presence of another protein (methionine aminopeptidase, or MetAP), rather than to the dynamics of TF, and model its TF-bound state. Based on these results, cryoENsemble is expected to find use for challenging heterogeneous cryo-EM maps for various biomolecular systems, especially those encompassing dynamic elements.

biophysics↗