Search bioRxiv⌕ Search

Biology subjects

Sanchez, I. E.

Publications and source records attributed to Sanchez, I. E..

3 recordsLinked to original sources

PhISCO: a simple method to infer phenotypes from protein sequences

Although protein sequences encode the information for folding and function, understanding their link is not an easy task. Unluckily, the prediction of how specific amino acids contribute to these features is still considerably impaired. Here, we developed PhISCO, Phenotype Inference from Sequence COmparisons, a simple algorithm that finds positions associated with any quantitative phenotype and predicts their values. From a few hundred sequences from four different protein families, we performed multiple sequence alignments and calculated per-position pairwise differences for both the sequence and the observed phenotypes. We found that from 3 to 10 positions, depending on the studied case, were enough to identify positions associated with the phenotypes and perform quantitative predictions of them. Here we show that these strong correlations can be found using individual positions while an improvement is achieved when the most correlated positions are jointly analyzed. Noteworthy, we performed phenotype predictions using a simple linear model that links per-position divergences and differences in observed phenotypes. We also show that although extremely simple, predictions are comparable to the state-of-art methodologies which, in most of the cases, are far more complex. All of the calculations are obtained at a very low information cost since the only input needed is a multiple sequence alignment of protein sequences with their associated quantitative phenotype. The diversity of the explored systems makes PhISCO a valuable tool to find sequence determinants of biological activity modulation and to predict various functional features for uncharacterized members of a protein family.

biophysics↗

Deamidation drives molecular aging of the SARS-CoV-2 spike receptor-binding motif

The spike is the main protein component of the SARS-CoV-2 virion surface. The spike receptor binding motif mediates recognition of the hACE2 receptor, a critical infection step, and is the preferential target for spike-neutralizing antibodies. Post-translational modifications of the spike receptor binding motif can modulate viral infectivity and immune response. We studied the spike protein in search for asparagine deamidation, a spontaneous event that leads to the appearance of aspartic and isoaspartic residues, affecting both the protein backbone and its charge. We used computational prediction and biochemical experiments to identify five deamidation hotspots in the SARS-CoV-2 spike. Similar deamidation hotspots are frequently found at the spike receptor-binding motifs of related sarbecoviruses, at positions that are mutated in emerging variants and in escape mutants from neutralizing antibodies. Asparagine residues 481 and 501 from the receptor-binding motif deamidate with a half-time of 16.5 and 123 days at 37 {degrees}C, respectively. This process is significantly slowed down at 4 {degrees}C, pointing at a strong dependence of spike molecular aging on the environmental conditions. Deamidation of the spike receptor-binding motif decreases the equilibrium constant for binding to the hACE2 receptor more than 3.5-fold. A model for deamidation of the full SARS-CoV-2 virion illustrates that deamidation of the spike receptor-binding motif leads to the accumulation in the virion surface of a chemically diverse spike population in a timescale of days. Our findings provide a mechanism for molecular aging of the spike, with significant consequences for understanding virus infectivity and vaccine development.

biochemistry↗

Conformational buffering underlies functional selection in intrinsically disordered protein regions

Many disordered proteins conserve essential functions in the face of extensive sequence variation. This makes it challenging to identify the forces responsible for functional selection. Viruses are robust model systems to investigate functional selection and they take advantage of protein disorder to acquire novel traits. Here, we combine structural and computational biophysics with evolutionary analysis to determine the molecular basis for functional selection in the intrinsically disordered adenovirus early gene 1A (E1A) protein. E1A competes with host factors to bind the retinoblastoma (Rb) protein, triggering early S-phase entry and disrupting normal cellular proliferation. We show that the ability to outcompete host factors depends on the picomolar binding affinity of E1A for Rb, which is driven by two binding motifs tethered by a hypervariable disordered linker. Binding affinity is determined by the spatial dimensions of the linker, which constrain the relative position of the two binding motifs. Despite substantial sequence variation across evolution, the linker dimensions are finely optimized through compensatory changes in amino acid sequence and sequence length, leading to conserved linker dimensions and maximal affinity. We refer to the mechanism that conserves spatial dimensions despite large-scale variations in sequence as conformational buffering. Conformational buffering explains how variable disordered proteins encode functions and could be a general mechanism for functional selection within disordered protein regions.

biophysics↗