Search bioRxiv⌕ Search

Biology subjects

Spiwok, V.

Publications and source records attributed to Spiwok, V..

3 recordsLinked to original sources

Minimal Amino Acid Alphabet for Protein Design

Proteins are built from 20 canonical amino acids. It is interesting to explore whether proteins can be formed from significantly reduced amino acid alphabets. Our bioinformatics survey of UniProt (more than 250 M sequences) revealed that proteins composed of reduced amino acid alphabets (< 10) are extremely rare among existing proteins. Next, we used computational protein design to design proteins composed of all 1,013 possible alphabets of 2-10 early amino acids (Ala, Asp, Glu, Gly, Ile, Leu, Pro, Ser, Thr, and Val). The length of all proteins was 100 amino acid residues. Small amino acid alphabets preferred simple helices or helix bundles. Larger amino acid alphabets allowed for the design of more complex structures. A protein composed of 8 amino acid types (Ala, Asp, Gly, Leu, Val, Ser, Thr, and Pro) was successfully experimentally verified. It adopts the {beta}-sheet-rich fibronectin type III domain architecture. Attempts to experimentally verify designs composed of 6 and 4 amino acid types were unsuccessful. We show by a computational experiment with an experimental validation that inverse folding models, namely ProteinMPNNsol, can stabilize a designed protein within the same eight-amino-acid alphabet. Our results show that globular proteins may have formed early in evolution. Furthermore, we show that it is possible to design proteins with interesting properties for biotechnology and synthetic biology.

bioinformatics↗

Design of proteins by parallel tempering in the sequence space

Design of new proteins is often formulated as an optimization task. An amino acid sequence is characterized by an energy, and this energy is sampled and minimized. Here, we use a parallel tempering algorithm to accelerate this task. A series of 100- or 200-residue proteins was designed using a modified Evolutionary Scale Modeling design module to maximize the confidence in structure prediction and globularity and minimize the surface hydrophobic residues. We show that parallel tempering is a viable alternative to Monte Carlo sampling and simulated annealing or related energy-based protein design methods, especially in the situation where a continuous flow of designed sequences is desired.

biochemistry↗

Free Energy Differences from Molecular Simulations: Exact Confidence Intervals from Transition Counts

Here we demonstrate a method to estimate the errors of free energy differences calculated by molecular simulations. The widths of the confidence intervals can be calculated solely from temperature and the number of transitions between states. Accuracy better than {+/-} 4.184 kJ/mol (1 kcal/mol) can be achieved by a simulation at 300 K with four forward and four reverse transitions. Markovianity of the process is a pre-requisite. For a two-state Markovian system, the confidence interval suggested below is exact (not only asymptotic or approximative), regardless the number of transitions. TOC Graphic O_FIG_DISPLAY_L [Figure 1] M_FIG_DISPLAY C_FIG_DISPLAY

biophysics↗