Search bioRxivSearch

Biology subjects

Claus O Wilke

Publications and source records attributed to Claus O Wilke.

7 recordsLinked to original sources

Sequence amplification via cell passaging creates spurious signals of positive adaptation in influenza H3N2 hemagglutinin

Clinical influenza A isolates are frequently not sequenced directly. Instead, a majority of these isolates (~70% in 2015) are first subjected to passaging for amplification, most commonly in non-human cell culture. Here, we find that this passaging leaves distinct signals of adaptation in the viral sequences, which can confound evolutionary analyses of the viral sequences. We find distinct patterns of adaptation to Madin-Darby (MDCK) and monkey cell culture absent from unpassaged hemagglutinin sequences. These patterns also dominate pooled datasets not separated by passaging type, and they increase in proportion to the number of passages performed. By contrast, MDCK-SIAT1 passaged sequences seem mostly (but not entirely) free of passaging adaptations. Contrary to previous studies, we find that using only internal branches of the influenza phylogenetic trees is insufficient to correct for passaging artifacts. These artifacts can only be safely avoided by excluding passaged sequences entirely from subsequent analysis. We conclude that future influenza evolutionary analyses should appropriately control for potentially confounding effects of passaging adaptations.

Evolutionary Biology

A comparison of one-rate and two-rate inference frameworks for site-specific dN/dS estimation

Two broad paradigms exist for inferring dN/dS, the ratio of nonsynonymous to synonymous substitution rates, from coding sequences: i) a one-rate approach, where dN/dS is represented with a single parameter, or ii) a two-rate approach, where dN and dS are estimated separately. The performances of these two approaches have been well-studied in the specific context of proper model specification, i.e. when the inference model matches the simulation model. By contrast, the relative performances of one-rate vs. two-rate parameterizations when applied to data generated according to a different mechanism remains unclear. Here, we compare the relative merits of one-rate and two-rate approaches in the specific context of model misspecification by simulating alignments with mutation-selection models rather than with dN/dS-based models. We find that one-rate frameworks generally infer more accurate dN/dS point estimates, even when dS varies among sites. In other words, modeling dS variation may substantially reduce accuracy of dN/dS point estimates. These results appear to depend on the selective constraint operating at a given site. In particular, for sites under strong purifying selection (dN/dS<~0.3), one-rate and two-rate models show comparable performances. However, one-rate models significantly outperform two-rate models for sites under moderate-to-weak purifying selection. We attribute this distinction to the fact that, for these more quickly evolving sites, a given substitution is more likely to be nonsynonymous than synonymous. The data will therefore be relatively enriched for nonsynonymous changes, and modeling dS contributes excessive noise to dN/dS estimates. We additionally find that high levels of divergence among sequences, rather than the number of sequences in the alignment, are more critical for obtaining precise point estimates.

Evolutionary Biology

Dissecting the roles of local packing density and longer-range effects in protein sequence evolution

What are the structural determinants of protein sequence evolution? A number of site-specific structural characteristics have been proposed, most of which are broadly related to either the density of contacts or the solvent accessibility of individual residues. Most importantly, there has been disagreement in the literature over the relative importance of solvent accessibility and local packing density for explaining site-specific sequence variability in proteins. We show here that this discussion has been confounded by the definition of local packing density. The most commonly used measures of local packing, such as the contact number and the weighted contact number, represent by definition the combined effects of local packing density and longer-range effects. As an alternative, we here propose a truly local measure of packing density around a single residue, based on the Voronoi cell volume. We show that the Voronoi cell volume, when calculated relative to the geometric center of amino-acid side chains, behaves nearly identically to the relative solvent accessibility, and both can explain, on average, approximately 34% of the site-specific variation in evolutionary rate in a data set of 209 enzymes. An additional 10% of variation can be explained by non-local effects that are captured in the weighted contact number. Consequently, evolutionary variation at a site is determined by the combined action of the immediate amino-acid neighbors of that site and of effects mediated by more distant amino acids. We conclude that instead of contrasting solvent accessibility and local packing density, future research should emphasize the relative importance of immediate contacts and longer-range effects on evolutionary variation.

Bioinformatics

Pyvolve: a flexible Python module for simulating sequences along phylogenies

We introduce Pyvolve, a flexible Python module for simulating genetic data along a phylogeny according to continuous-time Markov models of sequence evolution. Pyvolve is easily incorporated into Python bioinformatics pipelines, and it can simulate sequences according most standard models of nucleotide, amino-acid, and codon sequence evolution. All model parameters are fully customizable. Users can additionally specify custom evolutionary models, with custom rate matrices and/or states to evolve. This flexibility makes Pyvolve a convenient framework not only for simulating sequences under a wide variety of conditions, but also for developing and testing new evolutionary models. Moreover, Pyvolve includes several novel sequence simulation features, including a new rate matrix scaling algorithm and branch-length perturbations. Pyvolve is an open-source project freely available, along with a detailed user-manual and example scripts, under a FreeBSD license from http://github.com/sjspielman/pyvolve.

Bioinformatics

Geometric constraints dominate the antigenic evolution of influenza H3N2 hemagglutinin

We have carried out a comprehensive analysis of the determinants of human influenza A H3 hemagglutinin evolution, considering three distinct predictors of evolutionary variation at individual sites: solvent accessibility (as a proxy for protein fold stability and/or conservation), experimental epitope sites (as a proxy for host immune bias), and proximity to the receptor-binding region (as a proxy for protein function). We found that these three predictors individually explain approximately 15% of the variation in site-wise dN/dS. The solvent accessibility and proximity predictors were largely independent of each other, while the epitope sites were not. In combination, solvent accessibility and proximity explained 32% of the variation in dN/dS. Incorporating experimental epitope sites into the model added only an additional 2 percentage points. We also found that the historical H3 epitope sites, which date back to the 1980s and 1990s, showed only weak overlap with the latest experimental epitope data. Finally, sites with dN/dS > 1, i.e., the sites most likely driving seasonal immune escape, are not correctly predicted by either historical or experimental epitope sites, but only by proximity to the receptor-binding region. In summary, proximity to the receptor-binding region, and not host immune bias, seems to be the primary determinant of H3 evolution.\n\nAuthor summaryThe influenza virus is one of the most rapidly evolving human viruses. Every year, it accumulates mutations that allow it to evade the host immune response of previously infected individuals. Which sites in the virus genome allow this immune escape and the manner of escape is not entirely understood, but conventional wisdom states that specific \"immune epitope sites\" in the protein hemagglutinin are preferentially attacked by host antibodies and that these sites mutate to directly avoid host recognition; as a result, these sites are commonly targeted by vaccine development efforts. Here, we combine influenza hemagglutinin sequence data, protein structural information, experimental immune epitope data, and historical epitopes to demonstrate that neither the historical epitope groups nor epitopes based on experimental data are crucial for predicting the rate of influenza evolution. Instead, we find that a simple geometrical model works best: sites that are closest to the location where the virus binds the human receptor are the primary driver of hemagglutinin evolution. There are two possible explanations for this result. First, the existing historical and experimental epitope sites may not be the real antigenic sites in hemagglutinin. Second, alternatively, hemagglutinin antigenicity may not the primary driver of influenza evolution.

Evolutionary Biology

Intermediate Migration Yields Optimal Adaptation in Structured, Asexual Populations

Most evolving populations are subdivided into multiple subpopulations connected to each other by varying levels of gene flow. However, how population structure and gene flow (i.e., migration) affect adaptive evolution is not well understood. Here, we studied the impact of migration on asexually reproducing evolving computer programs (digital organisms). We found that digital organisms evolve the highest fitness values at intermediate migration rates, and we tested three hypotheses that could potentially explain this observation: (i) migration promotes passage through fitness valleys, (ii) migration increases genetic variation, and (iii) migration reduces clonal interference through a process called leapfrogging. We found that migration had no appreciable effect on the number of fitness valleys crossed and that genetic variation declined monotonously with increasing migration rates, instead of peaking at the optimal migration rate. However, the number of leapfrogging events, in which a superior beneficial mutation emerges on a genetic background that predates the previously best genotype in the population, did peak at the optimal migration rate. We thus conclude that in structured, asexual populations intermediate migration rates allow for optimal exploration of multiple, distinct fitness peaks, and thus yield the highest long-term adaptive success.

Evolutionary Biology

Predicting growth conditions from internal metabolic fluxes in an in-silico model of E. coli

A widely studied problem in systems biology is to predict bacterial phenotype from growth conditions, using mechanistic models such as flux balance analysis (FBA). However, the inverse prediction of growth conditions from phenotype is rarely considered. Here we develop a computational framework to carry out this inverse prediction on a computational model of bacterial metabolism. We use FBA to calculate bacterial phenotypes from growth conditions in E. coli, and then we assess how accurately we can predict the original growth conditions from the phenotypes. Prediction is carried out via regularized multinomial regression. Our analysis provides several important physiological and statistical insights. First, we show that by analyzing metabolic end products we can consistently predict growth conditions. Second, prediction is reliable even in the presence of small amounts of impurities. Third, flux through a relatively small number of reactions per growth source (~10) is sufficient for accurate prediction. Fourth, combining the predictions from two separate models, one trained only on carbon sources and one only on nitrogen sources, performs better than models trained to perform joint prediction. Finally, that separate predictions perform better than a more sophisticated joint prediction scheme suggests that carbon and nitrogen utilization pathways, despite jointly affecting cellular growth, may be fairly decoupled in terms of their dependence on specific assortments of molecular precursors.

Systems Biology