Search bioRxiv⌕ Search

Biology subjects

Salomon, J.

Publications and source records attributed to Salomon, J..

3 recordsLinked to original sources

Continuous B- to A- Transition in Protein-DNA Binding - How Well Is It Described by Current AMBER Force Fields?

When DNA interacts with a protein, its structure often undergoes significant conformational adaptation. Perhaps the most common is the transition from canonical B-DNA towards the A-DNA form, which is not a two-state, but rather a continuous transition. The A- and B-forms differ mainly in sugar pucker P (north/south) and glycosidic torsion {chi} (high-anti/anti). The combination of A-like P and B-like {chi} (and vice versa) represents the nature of the intermediate states lying between the pure A- and B- forms. In this work, we study how the A/B equilibrium and in particular the A/B intermediate states, which are known to be over-represented at protein-DNA interfaces, are modeled by current AMBER force fields. Eight protein-DNA complexes and their naked (unbound) DNAs were simulated with OL15 and bsc1 force fields as well as an experimental combination OL15{chi}OL3. We found that while the geometries of the A-like intermediate states in the molecular dynamics (MD) simulations agree well with the native X-ray geometries found in the protein-DNA complexes, their populations (stabilities) are significantly underestimated. Different force fields predict different propensities for A-like states growing in the order OL15 < bsc1 < OL15{chi}OL3, but the overall populations of the A-like form are too low in all of them. Interestingly, the force fields seem to predict the correct sequence-dependent A-form propensity, as they predict larger populations of the A-like form in naked (unbound) DNA in those steps that acquire A-like conformations in protein-DNA complexes. The instability of A-like geometries in current force fields may significantly alter the geometry of the simulated protein-DNA complex, destabilize the binding motif, and reduce the binding energy, suggesting that refinement is needed to improve description of protein-DNA interactions in AMBER force fields.

biophysics↗

NetSolP: predicting protein solubility in E. coli using language models

Solubility and expression levels of proteins can be a limiting factor for large-scale studies and industrial production. By determining the solubility and expression directly from the protein sequence, the success rate of wet-lab experiments can be increased. In this study, we focus on predicting the solubility and usability for purification of proteins expressed in Escherichia coli directly from the sequence. Our model NetSolP is based on deep learning protein language models called transformers and we show that it achieves state-of-the-art performance and improves extrapolation across datasets. As we find current methods are built on biased datasets, we curate existing datasets by using strict sequence-identity partitioning and ensure that there is minimal bias in the sequences. The predictor is available at https://services.healthtech.dtu.dk/service.php?NetSolP-1.0

bioinformatics↗

Deep protein representations enable recombinant protein expression prediction

A crucial process in the production of industrial enzymes is recombinant gene expression, which aims to induce enzyme overexpression of the genes in a host microbe. Current approaches for securing overexpression rely on molecular tools such as adjusting the recombinant expression vector, adjusting cultivation conditions, or performing codon optimizations. However, such strategies are time-consuming, and an alternative strategy would be to select genes for better compatibility with the recombinant host. Several methods for predicting soluble expression are available; however, they are all optimized for the expression host Escherichia coli and do not consider the possibility of an expressed protein not being soluble. We show that these tools are not suited for predicting expression potential in the industrially important host Bacillus subtilis. Instead, we build a B. subtilis-specific machine learning model for expressibility prediction. Given millions of unlabelled proteins and a small labeled dataset, we can successfully train such a predictive model. The unlabeled proteins provide a performance boost relative to using amino acid frequencies of the labeled proteins as input. On average, we obtain a modest performance of 0.64 area-under-the-curve (AUC) and 0.2 Matthews correlation coefficient (MCC). However, we find that this is sufficient for the prioritization of expression candidates for high-throughput studies. Moreover, the predicted class probabilities are correlated with expression levels. A number of features related to protein expression, including base frequencies and solubility, are captured by the model.

bioinformatics↗