Search bioRxivSearch

Biology subjects

Wilke, C. O.

Publications and source records attributed to Wilke, C. O..

12 recordsLinked to original sources

Evolutionary couplings detect side-chain interactions

Patterns of amino acid covariation in large protein sequence alignments can inform the prediction of de novo protein structures, binding interfaces, and mutational effects. While algorithms that detect these so-called evolutionary couplings between residues have proven useful for practical applications, less is known about how and why these methods perform so well, and what insights into biological processes can be gained from their application. Evolutionary coupling algorithms are commonly benchmarked by comparison to true structural contacts derived from solved protein structures. However, the methods used to determine true structural contacts are not standardized and different definitions of structural contacts may have important consequences for interpreting the results from evolutionary coupling analyses and understanding their overall utility. Here, we show that evolutionary coupling analyses are significantly more likely to identify structural contacts between side-chain atoms than between backbone atoms. We use both simulations and empirical analyses to highlight that purely backbone-based definitions of true residue-residue contacts (i.e., based on the distance between C atoms) may underestimate the accuracy of evolutionary coupling algorithms by as much as 40% and that a commonly used reference point (C{beta} atoms) underestimates the accuracy by 10-15%. These findings show that co-evolutionary outcomes differ according to which atoms participate in residue-residue interactions and suggest that accounting for different interaction types may lead to further improvements to contact-prediction methods.\n\nSignificance StatementEvolutionary couplings between residues within a protein can provide valuable information about protein structures, protein-protein interactions, and the mutability of individual residues. However, the mechanistic factors that determine whether two residues will co-evolve remains unknown. We show that structural proximity by itself is not sufficient for co-evolution to occur between residues. Rather, evolutionary couplings between residues are specifically governed by interactions between side-chain atoms. By contrast, intramolecular contacts between atoms in the protein backbone display only a weak signature of evolutionary coupling. These findings highlight that different types of stabilizing contacts exist within protein structures and that these types have a differential impact on the evolution of protein structures that should be considered in co-evolutionary applications.

biophysics

Theory of measurement for site-specific evolutionary rates in amino-acid sequences

In the field of molecular evolution, we commonly calculate site-specific evolutionary rates from alignments of amino-acid sequences. For example, catalytic residues in enzymes and interface regions in protein complexes can be inferred from observed relative rates. While numerous approaches exist to calculate amino-acid rates, it is not entirely clear what physical quantities the inferred rates represent and how these rates relate to the underlying fitness landscape of the evolving proteins. Further, amino-acid rates can be calculated in the context of different amino-acid exchangeability matrices, such as JTT, LG, or WAG, and again it is not well understood how the choice of the matrix influences the physical inter-pretation of the inferred rates. Here, we develop a theory of measurement for site-specific evolutionary rates, by analytically solving the maximum-likelihood equations for rate inference performed on sequences evolved under a mutation-selection model. We demonstrate that for realistic analysis settings the measurement process will recover the true expected rates of the mutation-selection model if rates are measured relative to a naive exchangeability matrix, in which all exchangeabilities are equal to 1/19. We also show that rate measurements using other matrices are quantitatively close but in general not mathematically equivalent. Our results demonstrate that insights obtained from phylogenetic-tree inference do not necessarily apply to rate inference, and best practices for the former may be deleterious for the latter.\n\nSignificance StatementMaximum likelihood inference is widely used to infer model parameters from sequence data in an evolutionary context. One major challenge in such inference procedures is the problem of having to identify the appropriate model used for inference. Model parameters usually are meaningful only to the extent that the model is appropriately specified and matches the process that generated the data. However, in practice, we dont know what process generated the data, and most models in actual use are misspecified. To circumvent this problem, we show here that we can employ maximum likelihood inference to make defined and meaningful measurements on arbitrary processes. Our approach uses misspecification as a deliberate strategy, and this strategy results in robust and meaningful parameter inference.

evolutionary biology

The many nuanced evolutionary consequences of duplicated genes

Gene duplication is seen as a major source of structural and functional divergence in genome evolution. Under the conventional models of sub- or neofunctionalizaton, functional changes arise in one of the duplicates after duplication. However, we suggest here that the presence of a duplicated gene can result in functional changes to its interacting partners. We explore this hypothesis by in-silico evolution of a heterodimer when one member of the interacting pair is duplicated. We examine how a range of selection pressures and protein structures leads to differential patterns of evolutionary divergence. We find that a surprising number of distinct evolutionary trajectories can be observed even in a simple three member system. Further, we observe that selection to correct dosage imbalance can affect the evolution of the initial function in several unexpected ways. For example, if a duplicate is under selective pressure to avoid binding its original binding partner, this can lead to changes in the binding interface of a non-duplicated interacting partner to exclude the duplicate. Hence, independent of the fate of the duplicate, its presence can impact how the original function operates. Additionally, we introduce a conceptual framework to describe how interacting partners cope with dosage imbalance after duplication. Contextualizing our results within this framework reveals that the evolutionary path taken by a duplicates interacting partners is highly stochastic in nature. Consequently, the fate of duplicate genes may not only be controlled by their own ability to accumulate mutations but also by how interacting partners cope with them.

evolutionary biology

Predicting bacterial growth conditions from mRNA and protein abundances

Cells respond to changing nutrient availability and external stresses by altering the expression of individual genes. Condition-specific gene expression patterns may provide a promising and low-cost route to quantifying the presence of various small molecules, toxins, or species-interactions in natural environments. However, whether gene expression signatures alone can predict individual environmental growth conditions remains an open question. Here, we used machine learning to predict 16 closely-related growth conditions using 155 datasets of E. coli transcript and protein abundances. We show that models are able to discriminate between different environmental features with a relatively high degree of accuracy. We observed a small but significant increase in model accuracy by combining transcriptome and proteome-level data, and we show that stationary phase conditions are typically more difficult to distinguish from one another than conditions under exponential growth. Nevertheless, with sufficient training data, gene expression measurements from a single species are capable of distinguishing between environmental conditions that are separated by a single environmental variable.

systems biology

Pinetree: a step-wise gene expression simulator with codon-specific translation rates

MotivationStochastic gene expression simulations often assume steady-state transcript levels, or they model transcription in more detail than translation. Moreover, they lack accessible programming interfaces, which limits their utility.\n\nResultsWe present Pinetree, a step-wise gene expression simulator with codon-specific translation rates. Pinetree models both transcription and translation in a stochastic framework with individual polymerase and ribosome-level detail. Written in C++ with a Python front-end, any user familiar with Python can specify a genome and simulate gene expression. Pinetree was designed to be efficient and scale to simulate large plasmids or viral genomes.\n\nAvailabilityPinetree is available on GitHub (https://github.com/benjaminjack/pinetree) and the Python Package Index (https://pypi.org/project/pinetree/).

systems biology

Combinatorial Approaches to Viral Attenuation

Attenuated viruses have numerous applications, in particular in the context of live viral vaccines. However, purposefully designing attenuated viruses remains challenging, in particular if the attenuation is meant to be resistant to rapid evolutionary recovery. Here we develop and analyze a new attenuation method, promoter ablation, using an established viral model, bacteriophage T7. Ablating promoters of the two most highly expressed T7 proteins (scaffold and capsid) led to major reductions in transcript abundance of the affected genes, with the effect of the double knockout approximately additive of the effects of single knockouts. Fitness reduction was moderate and also approximately additive; fitness recovery on extended adaptation was partial and did not restore the promoters. The fitness effect of promoter knockouts combined with a previously tested codon deoptimization of the capsid gene was less than additive, as anticipated from their competing mechanisms of action. In one design, the engineering created an unintended consequence that led to further attenuation, the effect of which was studied and understood in hindsight. Overall, the mechanisms and effects of genome engineering on attenuation behaved in a predictable manner. Therefore, this work suggests that the rational design of viral attenuation methods is becoming feasible.\n\nImportanceLive viral vaccines rely on attenuated viruses that can successfully infect their host but have reduced fitness or virulence. Such attenuated viruses were originally developed through trial- and-error, typically by adaptation of the wild-type virus to novel conditions. That method was haphazard, with no way of controlling the degree of attenuation, the number of attenuating mutations, or preventing evolutionary reversion. Synthetic biology now enables rational design and engineering of viral attenuation, but rational design must be informed by biological principles to achieve stable, quantitative attenuation. This work shows that in a model system for viral attenuation, bacteriophage T7, attenuation can be obtained from rational design principles, and multiple different attenuation approaches can be combined for enhanced overall effect.

microbiology

Selection removes Shine-Dalgarno-like sequences from within protein coding genes

The Shine-Dalgarno (SD) sequence motif facilitates translation initiation and is frequently found upstream of bacterial start codons. However, thousands of instances of this motif occur throughout the middle of protein coding genes in a typical bacterial genome. Here, we use comparative evolutionary analysis to test whether SD sequences located within genes are functionally constrained. We measure the conservation of SD sequences across Gammaproteobacteria, and find that they are significantly less conserved than expected. Further, the strongest SD sequences are the least conserved whereas we find evidence of conservation for the weakest possible SD sequences given amino acid constraints. Our findings indicate that most SD sequences within genes are likely to be deleterious and removed via selection. To illustrate the origin of these deleterious costs, we show that ATG start codons are significantly depleted downstream of SD sequences within genes, highlighting the potential for these sequences to promote erroneous translation initiation.

evolutionary biology

Limitation of alignment-free tools in total RNA-seq quantification

BackgroundAlignment-free RNA quantification tools have significantly increased the speed of RNA-seq analysis. However, it is unclear whether these state-of-the-art RNA-seq analysis pipelines can quantify small RNAs as accurately as they do with long RNAs in the context of total RNA quantification.\n\nResultWe comprehensively tested and compared four RNA-seq pipelines on the accuracies of gene quantification and fold-change estimation on a novel total RNA benchmarking dataset, in which small non-coding RNAs are highly represented along with other long RNAs. The four RNA-seq pipelines were of two commonly-used alignment-free pipelines and two variants of alignment-based pipelines. We found that all pipelines showed high accuracies for quantifying the expressions of long and highly-abundant genes. However, alignment-free pipelines showed systematically poorer performances in quantifying lowly-abundant and small RNAs.\n\nConclusionWe have shown that alignment-free and traditional alignment-based quantification methods performed similarly for common gene targets, such as protein-coding genes. However, we identified a potential pitfall in analyzing and quantifying lowly-expressed genes and small RNAs with alignment-free pipelines, especially when these small RNAs contain mutations.

bioinformatics

Beyond thermodynamic constraints: Evolutionary history shapes protein sequence variation

Biological evolution generates a surprising amount of site-specific variability in protein sequences. Yet attempts at modeling this process have been only moderately successful, and current models based on protein structural metrics explain, at best, 60% of the observed variation. Surprisingly, simple measures of protein structure, such as solvent accessibility, are often better predictors of site-specific variability than more complex models employing all-atom energy functions and detailed structural modeling. We suggest here that these more complex models perform poorly because they lack consideration of the evolutionary process that is in part captured by the simpler metrics. We compare protein sequences that are computationally designed to sequences that are computationally evolved using the same protein-design energy function and to homologous natural sequences. We find that by a wide variety of metrics, evolved sequences are much more similar to natural sequences than are designed sequences. In particular, designed sequences are too conserved on the protein surface relative to natural sequences whereas evolved sequences are not. Our results suggest that evolutionary simulation produces a realistic sampling of sequence space. By contrast, protein design--at least as currently implemented--does not. Existing energy functions seem to be sufficiently accurate to correctly describe the key thermodynamic constraints acting on protein sequences, but they need to be paired with realistic sampling schemes to generate realistic sequence alignments.

evolutionary biology

Reduced protein expression in a virus attenuated by codon deoptimization

The engineering of hundreds of synonymous codon changes into a viral genome appears to provide a general means of achieving attenuation. The mechanistic underpinnings of this approach remain enignmatic, however. Using quantitative proteomics and RNA sequencing, we explore the molecular basis of attenuation in a strain of bacteriophage T7 whose major capsid gene was engineered to carry 182 suboptimal codons. As expected, there was no evident effect of the recoding on transcription. Proteomic observations revealed that translation is halved for the recoded major capsid gene, and a smaller reduction applies to a few genes downstream, potentially caused by translational coupling. Viral burst size is also approximately halved, and the fitness drop accompanying attenuation is compatible with the reduced burst size. Overall, the fitness effect and molecular basis of attenuation by codon deoptimization are compatible with a relatively simple model of reduced translation of a few genes and a consequent diminished virion assembly. This mechanism is simpler than that operating in eukaryotic viruses.

systems biology

Accelerated simulation of evolutionary trajectories in origin-fixation models

We present an accelerated algorithm to forward-simulate origin--fixation models. Our algorithm requires on average only about two fitness evaluations per fixed mutation, whereas traditional algorithms require, per one fixed mutation, a number of fitness evaluations on the order of the effective population size Ne. Our accelerated algorithm yields the exact same steady state as the original algorithm but produces a different order of fixed mutations. By comparing several relevant evolutionary metrics, such as the distribution of fixed selection coefficients and the probability of reversion, we find that the two algorithms behave equivalently in many respects. However, the accelerated algorithm yields less variance in fixed selection coefficients. Notably, we are able to recover the expected amount of variance by rescaling population size, and we find a linear relationship between the rescaled population size and the population size used by the original algorithm. Considering the widespread usage of origin--fixation simulations across many areas of evolutionary biology, we introduce our accelerated algorithm as a useful tool for increasing the computational complexity of fitness functions without sacrificing much in terms of accuracy of the evolutionary simulation.

evolutionary biology

The E. coli molecular phenotype under different growth conditions

Modern systems biology requires extensive, carefully curated measurements of cellular components in response to different environmental conditions. While high-throughput methods have made transcriptomics and proteomics datasets widely accessible and relatively economical to generate, systematic measurements of both mRNA and protein abundances under a wide range of different conditions are still relatively rare. Here we present a detailed, genome-wide transcriptomics and proteomics dataset of E. coli grown under 34 different conditions. We manipulate concentrations of sodium and magnesium in the growth media, and we consider four different carbon sources glucose, gluconate, lactate, and glycerol. Moreover, samples are taken both in exponential and stationary phase, and we include two extensive time-courses, with multiple samples taken between 3 hours and 2 weeks. We find that exponential-phase samples systematically differ from stationary-phase samples, in particular at the level of mRNA. Regulatory responses to different carbon sources or salt stresses are more moderate, but we find numerous differentially expressed genes for growth on gluconate and under salt and magnesium stress. Our data set provides a rich resource for future computational modeling of E. coli gene regulation, transcription, and translation.

bioinformatics