Search bioRxivSearch

Biology subjects

Alexandre V Morozov

Publications and source records attributed to Alexandre V Morozov.

4 recordsLinked to original sources

Generalization of the Ewens sampling formula to arbitrary fitness landscapes

In considering evolution of transcribed regions, regulatory modules, and other genomic loci of interest, we are often faced with a situation in which the number of allelic states greatly exceeds the population size. In this limit, the population eventually adopts a steady state characterized by mutation-selection-drift balance. Although new alleles continue to be explored through mutation, the statistics of the population, and in particular the probabilities of seeing specific allelic configurations in samples taken from a population, do not change with time. In the absence of selection, probabilities of allelic configurations are given by the Ewens sampling formula, widely used in population genetics to detect deviations from neutrality. Here we develop an extension of this formula to arbitrary, possibly epistatic, fitness landscapes. Although our approach is general, we focus on the class of landscapes in which alleles are grouped into two, three, or several fitness states. This class of landscapes yields sampling probabilities that are computationally more tractable, and can form a basis for the inference of selection signatures from sequence data. We demonstrate that, for a sizeable range of mutation rates and selection coefficients, the steady-state allelic diversity is not neutral. Therefore, it may be used to infer selection coefficients, as well as other key evolutionary parameters, using high-throughput sequencing of evolving populations to collect data on locus polymorphisms. We also carry out numerical investigation of various approximations involved in deriving our sampling formulas, such as the infinite allele limit and the \"full connectivity\" assumption in which each allele can mutate into any other allele. We find that our theory remains sufficiently accurate even if these assumptions are relaxed. Thus, our framework establishes a theoretical foundation for inferring selection signatures from samples of sequences produced by evolution on epistatic fitness landscapes.

Genetics

Separating spandrels from phenotypic targets of selection in adaptive molecular evolution

There are many examples of adaptive molecular evolution in natural populations, but there is no existing method to verify which phenotypic changes were directly targeted by selection. The problem is that correlations between traits make it difficult to distinguish between direct and indirect selection. A phenotype is a direct target of selection when that trait in particular was shaped by selection to better perform a function. An indirect target of selection, also known as an evolutionary spandrel, is a phenotype that changes only because it is correlated with another trait under direct selection. Studies that mutate genes and examine the phenotypic consequences are increasingly common, and these experiments could estimate the mutational accessibility of the phenotypic changes that arise during an instance of adaptive molecular evolution. Under indirect selection, we expect phenotypes to evolve toward states that are more accessible by mutation. Deviation from this null expectation (evolution toward a phenotypic state rarely produced by mutation) would be compelling evidence of adaptation, and could be used to distinguish direct selection from indirect selection on correlated traits. To be practical, this molecular test of adaptation requires phenotypic differences that are caused by changes in a small number of genes. These kinds of genetically simple traits have been observed in many empirical studies of adaptive evolution. Here we describe how to use mutational accessibility to separate spandrels from direct targets of selection and thus verify adaptive hypotheses for phenotypes that evolve by adaptive molecular changes at one or a few genes.

Evolutionary Biology

A biophysical approach to predicting protein-DNA binding energetics

Sequence-specific interactions between proteins and DNA play a central role in DNA replication, repair, recombination, and control of gene expression. These interactions can be studied in vitro using microfluidics, protein-binding microarrays (PBMs), and other high-throughput techniques. Here we develop a biophysical approach to predicting protein-DNA binding specificities from high-throughput in vitro data. Our algorithm, called BindSter, accommodates multiple protein species competing for access to DNA and alternative binding modes of the same protein, while rigorously taking into account all sterically allowed configurations of DNA-bound particles. BindSter can be used with a hierarchy of protein-DNA interaction models of increasing complexity. We observe that the quality of BindSter predictions does not change significantly as some of the energy parameters vary over a sizable range. To take this degeneracy into account, we have developed a graphical representation of parameter uncertainties, called IntervalLogo. We find that our simplest model, in which each nucleotide in the binding site is treated independently, performs better than previous biophysical approaches. The extensions of this model, in which contributions of longer words are also considered, result in further improvements, underscoring the importance of higherorder effects in protein-DNA energetics. In contrast, we find little evidence for multiple binding modes for the transcription factors (TFs) in our dataset. Furthermore, there is limited consistency in predictions for the same TF utilizing microfluidics and PBM experimental platforms.

Biophysics

Protein folding and binding can emerge as evolutionary spandrels through structural coupling

Binding interactions between proteins and other molecules mediate numerous cellular processes, including metabolism, signaling, and regulation of gene expression. These interactions evolve in response to changes in the protein's chemical or physical environment (such as the addition of an antibiotic), or when genes duplicate and diverge. Several recent studies have shown the importance of folding stability in constraining protein evolution. Here we investigate how structural coupling between protein folding and binding -- the fact that most proteins can only bind their targets when folded -- gives rise to evolutionary coupling between the traits of folding stability and binding strength. Using biophysical and evolutionary modeling, we show how these protein traits can emerge as evolutionary "spandrels" even if they do not confer an intrinsic fitness advantage. In particular, proteins can evolve strong binding interactions that have no functional role but merely serve to stabilize the protein if misfolding is deleterious. Furthermore, such proteins may have divergent fates, evolving to bind or not bind their targets depending on random mutation events. These observations may explain the abundance of apparently nonfunctional interactions among proteins observed in high-throughput assays. In contrast, for proteins with both functional binding and deleterious misfolding, evolution may be highly predictable at the level of biophysical traits: adaptive paths are tightly constrained to first gain extra folding stability and then partially lose it as the new binding function is developed. These findings have important consequences for our understanding of fundamental evolutionary principles of both natural and engineered proteins.

Evolutionary Biology