Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Evolutionary Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

A novel Bayesian method for inferring and interpreting the dynamics of adaptive landscapes from phylogenetic comparative data

Our understanding of macroevolutionary patterns of adaptive evolution has greatly increased with the advent of large-scale phylogenetic comparative methods. Widely used Ornstein-Uhlenbeck (OU) models can describe an adaptive process of divergence and selection. However, inference of the dynamics of adaptive landscapes from comparative data is complicated by interpretational difficulties, lack of identifiability among parameter values and the common requirement that adaptive hypotheses must be assigned a priori. Here we develop a reversible-jump Bayesian method of fitting multi-optima OU models to phylogenetic comparative data that estimates the placement and magnitude of adaptive shifts directly from the data. We show how biologically informed hypotheses can be tested against this inferred posterior of shift locations using Bayes Factors to establish whether our a priori models adequately describe the dynamics of adaptive peak shifts. Furthermore, we show how the inclusion of informative priors can be used to restrict models to biologically realistic parameter space and test particular biological interpretations of evolutionary models. We argue that Bayesian model-fitting of OU models to comparative data provides a framework for integrating of multiple sources of biological data--such as microevolutionary estimates of selection parameters and paleontological timeseries--allowing inference of adaptive landscape dynamics with explicit, process-based biological interpretations.

Evolutionary Biology

Mechanical interactions in bacterial colonies and the surfing probability of beneficial mutations

Bacterial conglomerates such as biofilms and microcolonies are ubiquitous in nature and play an important role in industry and medicine. In contrast to well-mixed, diluted cultures routinely used in microbial research, bacteria in a microcolonv interact mechanically with one another and with the substrate to which they are attached. Despite their ubiquity, little is known about the role of such mechanical interactions on growth and biological evolution of microbial populations. Here we use a computer model of a microbial colony of rod-shaped cells to investigate how physical interactions between cells determine their motion in the colony, this affects biological evolution. We show that the probability that a faster-growing mutant \"surfs\" at the colonys frontier and creates a macroscopic sector depends on physical properties of cells (shape, elasticity, friction). Although all these factors contribute to the surfing probability in seemingly different ways, they all ultimately exhibit their effects by altering the roughness of the expanding frontier of the colony and the orientation of cells. Our predictions are confirmed by experiments in which we measure the surfing probability for colonies of different front roughness. Our results show that physical interactions between bacterial cells play an important role in biological evolution of new traits, and suggest that these interaction may be relevant to processes such as de novo evolution of antibiotic resistance.

evolutionary biology

The genealogical sorting index and species delimitation

The Genealogical Sorting Index (gsi) has been widely used in species-delimitation studies, where it is usually interpreted as a measure of the degree to which each of several predefined groups of specimens display a pattern of divergent evolution in a phylogenetic tree. Here we show that the gsi value obtained for a given group is highly dependent on the structure of the tree outside of the group of interest. By calculating the gsi from simulated datasets we demonstrate this dependence undermines some of desirable properties of the statistic. We also review the use of the gsi delimitation studies, and show that the gsi has typically been used under scenarios in which it is expected to produce large and statistically significant results for samples that are not divergent from all other populations and thus should not be considered species. Our proposed solution to this problem performs better than the gsi in under these conditions. Nevertheless, we show that our modified approach can produce positive results for populations that are connected by substantial levels of gene flow, and are thus unlikely to represent distinct species. We stress that the properties of gsi made clear in this manuscript must be taken into account if the statistic is used in species-delimitation studies. More generally, we argue that the results of genetic species-delimitation methods need to be interpreted in the light the biological and ecological setting of a study, and not treated as the final test applied to hypotheses generated by other data.

Evolutionary Biology

Population Genetics Based Phylogenetics Under Stabilizing Selection for an Optimal Amino Acid Sequence: A Nested Modeling Approach

We present a new phylogenetic approach SelAC (Selection on Amino acids and Codons), whose substitution rates are based on a nested model linking protein expression to population genetics. Unlike simpler codon models which assume a single substitution matrix for all sites, our model more realistically represents the evolution of protein coding DNA under the assumption of consistent, stabilizing selection using cost-benefit approach. This cost-benefit approach allows us generate a set of 20 optimal amino acid specific matrix families using just a handful of parameters and naturally links the strength of stabilizing selection to protein synthesis levels, which we can estimate. Using a yeast dataset of 100 orthologs for 6 taxa, we find SelAC fits the data much better than popular models by 104-105 AICc units. Our results indicate there is great potential for more accurate inference of phylogenetic trees and branch lengths from already existing data through the use of nested, mechanistic models. Additional parameters estimated by SelAC indicate that a large amount of non-phylogenetic, but biologically meaningful, information can be inferred from exisiting data. For example, SelAC prediction of gene specific protein synthesis rates correlates well with both empirical (r=0.33-0.48) and other theoretical predictions (r=0.45-0.64) for multiple yeast species. SelAC also provides estimates of the optimal amino acid at each site. Finally, because SelAC is a nested approach based on clearly stated biological assumptions, future modifications, such as including shifts in the optimal amino acid sequence within or across lineages, are possible.

evolutionary biology

Gene tree discordance causes apparent substitution rate variation

Substitution rates are known to be variable among genes, chromosomes, species, and lineages due to multifarious biological processes. Here we consider another source of substitution rate variation due to a technical bias associated with gene tree discordance, which has been found to be rampant in genome-wide datasets, often due to incomplete lineage sorting (ILS). This apparent substitution rate variation is caused when substitutions that occur on discordant gene trees are analyzed in the context of a single, fixed species tree. Such substitutions have to be resolved by proposing multiple substitutions on the species tree, and we therefore refer to this phenomenon as \"SPILS\" (Substitutions Produced by Incomplete Lineage Sorting). We use simulations to demonstrate that SPILS has a larger effect with increasing levels of ILS, and on trees with larger numbers of taxa. Specific branches of the species trees are consistently, but erroneously, inferred to be longer or shorter, and we show that these branches can be predicted based on discordant tree topologies. Moreover, we observe that fixing a species tree topology when performing tests of positive selection increases the false positive rate, particularly for genes whose discordant topologies are most affected by SPILS. Finally, we use data from multiple Drosophila species to show that SPILS can be detected in nature. While the effects of SPILS are modest per gene, it has the potential to affect substitution rate variation whenever high levels of ILS are present, particularly in rapid radiations. The problems outlined here have implications for character mapping of any type of trait, and for any biological process that causes discordance. We discuss possible solutions to these problems, and areas in which they are likely to have caused faulty inferences of convergence and accelerated evolution.

Evolutionary Biology

RAD sequencing and a hybrid Antarctic fur seal genome assembly reveal rapidly decaying linkage disequilibrium, global population structure and evidence for inbreeding

Recent advances in high throughput sequencing have transformed the study of wild organisms by facilitating the generation of high quality genome assemblies and dense genetic marker datasets. These resources have the potential to significantly advance our understanding of diverse phenomena at the level of species, populations and individuals, ranging from patterns of synteny through rates of linkage disequilibrium (LD) decay and population structure to individual inbreeding. Consequently, we used PacBio sequencing to refine an existing Antarctic fur seal (Arctocephalus gazella) genome assembly and genotyped 83 individuals from six populations using restriction site associated DNA (RAD) sequencing. The resulting hybrid genome comprised 6,169 scaffolds with an N50 of 6.21 Mb and provided clear evidence for the conservation of large chromosomal segments between the fur seal and dog (Canis lupus familiaris). Focusing on the most extensively sampled population of South Georgia, we found that LD decayed rapidly, reaching the background level of r2 = 0.09 by around 26 kb, consistent with other vertebrates but at odds with the notion that fur seals experienced a strong historical bottleneck. We also found evidence for population structuring, with four main Antarctic island groups being resolved. Finally, appreciable variance in individual inbreeding could be detected, reflecting the strong polygyny and site fidelity of the species. Overall, our study contributes important resources for future genomic studies of fur seals and other pinnipeds while also providing a clear example of how high throughput sequencing can generate diverse biological insights at multiple levels of organisation.

evolutionary biology

A Likelihood-Free Inference Framework for Population Genetic Data using Exchangeable Neural Networks

An explosion of high-throughput DNA sequencing in the past decade has led to a surge of interest in population-scale inference with whole-genome data. Recent work in population genetics has centered on designing inference methods for relatively simple model classes, and few scalable general-purpose inference techniques exist for more realistic, complex models. To achieve this, two inferential challenges need to be addressed: (1) population data are exchangeable, calling for methods that efficiently exploit the symmetries of the data, and (2) computing likelihoods is intractable as it requires integrating over a set of correlated, extremely high-dimensional latent variables. These challenges are traditionally tackled by likelihood-free methods that use scientific simulators to generate datasets and reduce them to hand-designed, permutation-invariant summary statistics, often leading to inaccurate inference. In this work, we develop an exchangeable neural network that performs summary statistic-free, likelihood-free inference. Our frame-work can be applied in a black-box fashion across a variety of simulation-based tasks, both within and outside biology. We demonstrate the power of our approach on the recombination hotspot testing problem, outperforming the state-of-the-art.

evolutionary biology

Determinants of genetic structure of the Sub-Saharan parasitic wasp Cotesia sesamiae

Parasitoid life style represents one of the most diversified life history strategies on earth. There are however very few studies on the variables associated with intraspecific diversity of parasitoid insects, especially regarding the relationship with spatial, biotic and abiotic ecological factors. Cotesia sesamiae is a Sub-Saharan stenophagous parasitic wasp that parasitizes several African stemborer species with variable developmental success. The different host-specialized populations are infected with different strains of Wolbachia, an endosymbiotic bacterium widespread in arthropods that is known for impacting life history traits notably reproduction, and consequently species distribution. In this study, first we analyzed the genetic structure of C. sesamiae across Sub-Saharan Africa, using 8 microsatellite markers, and 3 clustering software. We identified five major population clusters across Sub-Saharan Africa, which probably originated in East African Rift region and expanded throughout Africa in relation to host genus and abiotic factors such as climatic classifications. Using laboratory lines, we estimated the incompatibility between the different strains of Wolbachia infecting C. sesamiae. We observed an incompatibility between Wolbachia strains was asymmetric; expressed in one direction only. Based on these results, we assessed the relationships between direction of gene flow and Wolbachia infections in the genetic clusters. We found that Wolbachia-induced reproductive incompatibility was less influential than host specialization in the genetic structure. Both Wolbachia and host were more influential than geography and current climatic conditions. These results are discussed in the context of African biogeography, and co-evolution between Wolbachia, virus parasitoid and host, in the perspective of improving biological control efficiency through a better knowledge of the biodiversity of biological control agents.

evolutionary biology

A note on measuring natural selection on principal component scores

Measuring natural selection through the use of multiple regression has transformed our understanding of selection, although the methods used remain sensitive to the effects of multicollinearity due to highly correlated traits. While measuring selection on principal component scores is an apparent solution to this challenge, this approach has been heavily criticized due to difficulties in interpretation and relating PC axes back to the original traits. We describe and illustrate how to transform selection gradients for PC scores back into selection gradients for the original traits, addressing issues of multicollinearity and biological interpretation. We demonstrate this approach with empirical data and examples from the literature, highlighting how selection estimates for PC scores can be interpreted while reducing the consequences of multicollinearity.

evolutionary biology

Coalescent Processes With Skewed Offspring Distributions And Non-Equilibrium Demography

Non-equilibrium demography impacts coalescent genealogies leaving detectable, well-studied signatures of variation. However, similar genomic footprints are also expected under models of large reproductive skew, posing a serious problem when trying to make inference. Furthermore, current approaches consider only one of the two processes at a time, neglecting any genomic signal that could arise from their simultaneous effects, preventing the possibility of jointly inferring parameters relating to both offspring distribution and population history. Here, we develop an extended Moran model with exponential population growth, and demonstrate that the underlying ancestral process converges to a time-inhomogeneous psi-coalescent. However, by applying a non-linear change of time scale - analogous to the Kingman coalescent - we find that the ancestral process can be rescaled to its time-homogeneous analogue, allowing the process to be simulated quickly and efficiently. Furthermore, we derive analytical expressions for the expected site-frequency spectrum under the time-inhomogeneous psi-coalescent and develop an approximate-likelihood framework for the joint estimation of the coalescent and growth parameters. By means of extensive simulation, we demonstrate that both can be estimated accurately given linkage equilibrium, while linkage disequilibrium systematically biases growth rate estimates. In addition, not accounting for demography can lead to serious biases in the inferred coalescent model, with broad implications for genomic studies ranging from ecology to conservation biology. Finally, we use our method to analyze sequence data from Japanese sardine populations and find evidence of high variation in individual reproductive success, but few signs of a recent demographic expansion.

evolutionary biology

On the concept of biological function, junk DNA and the gospels of ENCODE and Graur et al.

In a recent article entitled \"On the immortality of television sets: \"function\" in the human genome according to the evolution-free gospel of ENCODE\", Graur et al. dismantle ENCODEs evidence and conclusion that 80% of the human genome is functional. However, the article by Graur et al. contains assumptions and statements that are questionable. Primarily, the authors limit their evaluation of DNAs biological functions to informational roles, sidestepping putative non-informational functions. Here, I bring forward an old hypothesis on the evolution of genome size and on the role of so called junk DNA (jDNA), which might explain C-value enigma. According to this hypothesis, the jDNA functions as a defense mechanism against insertion mutagenesis by endogenous and exogenous inserting elements such as retroviruses, thereby protecting informational DNA sequences from inactivation or alteration of their expression. Notably, this model couples the mechanisms and the selective forces responsible for the origin of jDNA with its putative protective biological function, which represents a classic example of fighting fire with fire. One of the key tenets of this theory is that in humans and many other species, jDNAs serves as a protective mechanism against insertional oncogenic transformation. As an adaptive defense mechanism, the amount of protective DNA varies from one species to another based on the rate of its origin, insertional mutagenesis activity, and evolutionary constraints on genome size.

Evolutionary Biology

Genome-wide patterns of local adaptation in Drosophila melanogaster: adding intra European variability to the map

Signatures of spatially varying selection have been investigated both at the genomic and transcriptomic level in several organisms. In Drosophila melanogaster, the majority of these studies have analyzed North American and Australian populations, leading to the identification of several loci and traits under selection. However, populations in these two continents showed evidence of admixture that likely contributed to the observed population differentiation patterns. Thus, disentangling demography from selection is challenging when analyzing these populations. European populations could be a suitable system to identify loci under spatially varying selection provided that no recent admixture from African populations would have occurred. In this work, we individually sequence the genome of 42 European strains collected in populations from contrasting environments: Stockholm (Sweden), and Castellana Grotte, (Southern Italy). We found low levels of population structure and no evidence of recent African admixture in these two populations. We thus look for patterns of spatially varying selection affecting individual genes and gene sets. Besides single nucleotide polymorphisms, we also investigate the role of transposable elements in local adaptation. We concluded that European populations are a good dataset to identify loci under spatially varying selection. The analysis of the two populations sequenced in this work in the context of all the available D. melanogaster data allowed us to pinpoint genes and biological processes relevant for local adaptation. Identifying and analyzing populations with low levels of population structure and admixture should help to disentangle selective from non-selective forces underlying patterns of population differentiation in other species as well.

evolutionary biology

Modeling the growth of organisms validates a general relation between metabolic costs and natural selection

Metabolism and evolution are closely connected: if a mutation incurs extra energetic costs for an organism, there is a baseline selective disadvantage that may or may not be compensated for by other adaptive effects. A long-standing, but to date unproven, hypothesis is that this disadvantage is equal to the fractional cost relative to the total resting metabolic expenditure. This hypothesis has found a recent resurgence as a powerful tool for quantitatively understanding the strength of selection among different classes of organisms. Our work explores the validity of the hypothesis from first principles through a generalized metabolic growth model, versions of which have been successful in describing organismal growth from single cells to higher animals. We build a mathematical framework to calculate how perturbations in maintenance and synthesis costs translate into contributions to the selection coefficient, a measure of relative fitness. This allows us to show that the hypothesis is an approximation to the actual baseline selection coefficient. Moreover we can directly derive the correct prefactor in its functional form, as well as analytical bounds on the accuracy of the hypothesis for any given realization of the model. We illustrate our general framework using a special case of the growth model, which we show provides a quantitative description of overall metabolic synthesis and maintenance expenditures in data collected from a wide array of unicellular organisms (both prokaryotes and eukaryotes). In all these cases we demonstrate that the hypothesis is an excellent approximation, allowing estimates of baseline selection coefficients to within 15% of their actual values. Even in a broader biological parameter range, covering growth data from multicellular organisms, the hypothesis continues to work well, always within an order of magnitude of the correct result. Our work thus justifies its use as a versatile tool, setting the stage for its wider deployment.

evolutionary biology

Measuring Selection Across HIV Gag: Combining Physico-Chemistry and Population Genetics

We present physico-chemical based model grounded in population genetics. Our model predicts the stationary probability of observing an amino acid residue at a given site. Its predictions are based on the physico-chemical properties of the inferred optimal residue at that site and the sensitivity of the proteins functionality to deviation from the physico-chemical optimum at that site. We contextualize our physico-chemical model by comparing our model fit and parameters it to the more general, but less biologically meaningful entropy based metric: site sensitivity or 1/E. We show mathematically that our physico-chemical model is a more restricted form of the entropy model and how 1/E is proportional to the log-likelihood of a parameter-wise saturated model. Next, we fit both our physico-chemical and entropy models to sequences for subtype Cs Gag poly-protein in the LANL HIV database. Comparing our models site sensitivity parameters G' to 1/E we find they are highly correlated. We also compare the ability of G', 1/E, and other indirect measures of HIV fitness to empirical in vitro and in vivo measures. We find G' does a slightly better job predicting empirical fitness measures of in vivo viral escape time and in vitro spreading rates. While our predictive gain is modest, our model can be modified to test more complex or alternative biological hypotheses. More generally, because of its explicit biological formulation, our model can be easily extended to test for stabilizing vs. diversifying selection. We conjecture that our model could also be extended include epistasis in a more realistic manner than Ising models, while requiring many fewer parameters than Potts models.

evolutionary biology

A framework for estimating the effects of sequential reproductive barriers: implementation using Bayesian models with field data from cryptic species

Determining how reproductive barriers modulate gene flow between populations represents a major step towards understanding the factors shaping the course of speciation. Although many indices quantifying reproductive isolation (RI) have been proposed, they do not permit the quantification of cross direction-specific RI under varying species frequencies and over arbitrary sequences of barriers. Furthermore, techniques quantifying associated uncertainties are lacking, and statistical methods unrelated to biological process are still preferred for obtaining confidence intervals and p-values. To address these shortcomings, we provide new RI indices that model changes in gene flow for both directions of hybridization, and we implement them in a Bayesian model. We use this model to quantify RI between two species of the psyllid Cacopsylla pruni based on field genotypic data for mating individuals, inseminated spermatophores and progeny. The results showed that pre-insemination isolation was strong, mildly asymmetric and undistinguishably different between study sites despite large differences in species frequencies; that post-insemination isolation strongly affected the more common hybrid type; and that cumulative isolation was close to complete. In the light of these results, we discuss how these developments can strengthen comparative RI studies.\n\nAuthor contributionsJP and NS initiated the study and obtained biological data. JP and DRJP developed the porosity-based approach. DRJP conceived the Bayesian implementation and code. JP, DRJP and NS wrote the manuscript.\n\nData availabilityMitochondrial sequence data will be available at Genbank, source code is available at xxx.

evolutionary biology

Adaptive evolution by spontaneous domain fusion and protein relocalisation

Knowledge of adaptive processes encompasses understanding of the emergence of new genes. Computational analyses of genomes suggest that new genes can arise by domain swapping, however, empirical evidence has been lacking. Here we describe a set of nine independent deletion mutations that arose during the course of selection experiments with the bacterium Pseudomonas fluorescens in which the membrane-spanning domain of a fatty acid desaturase became translationally fused to a cytosolic di-guanylate cyclase (DGC) generating an adaptive phenotype. Detailed genetic analysis of one chimeric fusion protein showed that the DGC domain had become membrane-localised resulting in a new biological function. The relative ease by which this new gene arose along with its profound functional and regulatory effects provides a glimpse of mutational events and their consequences that are likely to play a significant role in the evolution of new genes.

evolutionary biology

Targeted sequencing of venom genes from cone snail genomes reveals coupling between dietary breadth and conotoxin diversity

Although venomous taxa provide an attractive system to study the genetic basis of adaptation and speciation, the slow pace of toxin gene discovery through traditional laboratory techniques (e.g., cDNA cloning) have limited their utility in the study of ecology and evolution. Here, we applied targeted sequencing techniques to selectively recover venom gene superfamilies and non-toxin loci from the genomes of 32 species of cone snails (family, Conidae), a hyper diverse group of carnivorous marine gastropods that capture their prey using a cocktail of neurotoxic proteins (conotoxins). We were able to successfully recover conotoxin gene superfamilies across all species sequenced in this study with high confidence (> 100X coverage). We found that conotoxin gene superfamilies are composed of 1-6 exons and adjacent noncoding regions are not enriched for simple repetitive elements. Additionally, we provided further evidence for several genetic factors shaping venom composition in cone snails, including positive selection, extensive gene turnover, expression regulation, and potentially, presence-absence variation. Using comparative phylogenetic methods, we found that while diet specificity did not predict patterns of conotoxin gene superfamily size evolution, dietary breadth was positively correlated with total conotoxin gene diversity. These results continue to emphasize the importance of dietary breadth in shaping venom evolution, an underappreciated ecological correlate in venom biology. Finally, the targeted sequencing technique demonstrated here has the potential to radically increase the pace at which venom gene families are sequenced and studied, reshaping our ability to understand the impact of genetic changes on ecologically relevant phenotypes and subsequent diversification.

evolutionary biology

An amplicon-based sequencing framework for accurately measuring intrahost virus diversity using PrimalSeq and iVar

How viruses evolve within hosts can dictate infection outcomes; however, reconstructing this process is challenging. We evaluated our multiplexed amplicon approach - PrimalSeq - to demonstrate how virus concentration, sequencing coverage, primer mismatches, and replicates influence the accuracy of measuring intrahost virus diversity. We developed an experimental protocol and computational tool (iVar) for using PrimalSeq to measure virus diversity using Illumina and compared the results to Oxford Nanopore sequencing. We demonstrate the utility of PrimalSeq by measuring Zika and West Nile virus diversity from varied sample types and show that the accumulation of genetic diversity is influenced by experimental and biological systems.

evolutionary biology