Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Evolutionary Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 811 records · Page 45Linked to original sources

A statistical approach to genome size evolution: Observations and explanations

Genome size evolution is a fundamental problem in molecular evolution. Statistical analysis of genome sizes brings new insight into the evolution of genome size. Although the variation of genome sizes is complicated, it is indicated that the genome size evolution can be explained more clearly at taxon level than at species level. I find that the genome size distribution for species in a taxon fits log-normal distribution. And I find a relationship between the phylogeny of life and the statistical features of genome size distributions among taxa. I observed different statistical features of genome size distributions between animal taxa and plant taxa. A log-normal stochastic process model is developed to simulate the genome size evolution. The simulation results on the log-normal distributions of genome sizes and their statistical features agree with the observations.

Evolutionary Biology

Population size and the length of the chromosome blocks identical by descent over generations

In all populations, as the time runs, crossovers break apart ancestor haplotypes, forming smaller blocks at each generation. Some blocks, and eventually all of them, become identical by descent because of the genetic drift. We have in this paper developed and benchmarked a theoretical prediction of the mean length of such blocks and used it to study a simple population model assuming panmixia, no selfing and drift as the only evolutionary pressure. Besides, we have on the one hand derived, for any user defined error threshold, the range of the parameters this prediction is reliable for, and on the other hand shown that the mean length remains constant over time in ideally large populations.

Evolutionary Biology

An in silico comparison of reduced-representation and sequence-capture protocols for phylogenomics

In the age of genome-scale DNA sequencing, choice of molecular marker arguably remains an important decision in planning a phylogenetic study. Using published genomes from 23 primate species, we make a standardized comparison of four of the most frequently used protocols in phylogenomics, viz., targeted sequence-enrichment using ultraconserved element and exon-capture probes, and reduced genomic representation using restriction-site-associated DNA sequencing (RADseq and ddRAD-seq). Here we present a procedure to perform in silico extractions from genomes and create directly comparable datasets for each class of marker. We then compare these datasets in terms of both phylogenetic resolution and ability to consistently and precisely estimate clade ages using fossil-calibrated molecular-clock models. Furthermore, we were also able to directly compare these results to previously published datasets from Sanger-sequenced nuclear exons and mitochondrial genomes under the same analytical conditions. Our results show--although with the exception of the mitochondrial genome and ddRADseq datasets--that for uncontroversial nodes all data classes performed equally well, i.e. they recovered the same well supported topology. However, for one difficult-to-resolve node comprising a rapid diversification (subfamilial relationships among the Cebidae), we report well supported but conflicting topologies among the marker classes, likely the result of mismodelling of gene tree heterogeneity. Likewise, clade age estimates showed consistent discrepancies between datasets; for recent nodes, clade ages estimated by nuclear exon datasets were younger than those of the UCE, RAD and mitochondrial data, but vice versa for the deepest nodes in the primate phylogeny. This effect can be explained by temporal differences in phylogenetic informativeness and choice of clock model used. Finally, we conclude by emphasizing that while huge numbers of loci are probably not required for uncontroversial phylogenetic questions--for which practical considerations such as cost and ease of data generation/sharing/aggregating therefore become increasingly important--accurately modelling heterogeneous data remains as relevant as ever for the more recalcitrant problems.

Evolutionary Biology

Virility does not Imply Immensity: Testis size, Accessory Gland Size and Ejaculate depletion pattern do not Evolve in Response to Experimental Manipulation of Sex Ratio in Drosophila melanogaster

Introduction Introduction Materials and Methods Dissections and Measurements Results Discussion References Females in many species mate more than once and store sperm from more than one male. This leads to post-copulatory competition where sperm from different males compete to fertilize the limited number of eggs produced by the female-typically called Sperm competition (Parker, 1970a) (Wedell et al., 2002). Sperm competition and the resulting post-copulatory selection can significantly alter male reproductive behavior (Simmons et al., 1993) (Cook & Wedell, 1996) (Gage and Barnard, 1996) (Wedell & Cook 1999a,b) (Bretman et al., 2009) (Bretman et al., 2010) ( ...

Evolutionary Biology

Rapid evolution of the inter-sexual genetic correlation for fitness in Drosophila melanogaster

Sexual antagonism (SA) arises when male and female phenotypes are under opposing selection, yet genetically correlated. Until resolved, antagonism limits evolution towards optimal sex-specific phenotypes. Despite its importance for sex-specific adaptation and existing theory, the dynamics of SA resolution are not well understood empirically. Here, we present data from Drosophila melanogaster, compatible with a resolution of SA. We compared two independent replicates of the LHM population in which SA had previously been described. Both had been maintained under identical, controlled conditions, and separated for <250 generations. Although heritabilities of male and female fitness were similar, the inter-sexual genetic correlation differed significantly, being negative in one replicate (indicating SA) but close to zero in the other. Using population sequencing, we show that phenotypic differences were associated with population divergence in allele frequencies at non-random loci across the genome. Large frequency changes were more prevalent in the population without SA and were enriched at loci mapping to genes previously shown to have a sexually antagonistic relationships between expression and fitness. Our data suggest that rapid evolution towards SA resolution has occurred in one of the populations and open avenues towards studying the genetics of SA and its resolution.

Evolutionary Biology

Analysis of the optimality of the Standard Genetic Code.

Many theories have been proposed attempting to explain the origin of the genetic code. While strong reasons remain to believe that the genetic code evolved as a frozen accident, at least for the first few amino acids, other theories remain viable. In this work, we test the optimality of the standard genetic code against approximately 17 million genetic codes, and locate 18 which outperform the standard genetic code at the following three criteria: (a) robustness to point mutation; (b) robustness to frameshift mutation; and (c) ability to encode additional information in the coding region. We use a genetic algorithm to generate and score codes from different parts of the associated landscape, and are, as a result presumably more representative of the entire landscape. Our results show that while the genetic code is sub-optimal for robustness to frameshift mutation and the ability to encode additional information in the coding region, it is very strongly selected for robustness to point mutation. This coupled with the observation that the different performance indicator scores for a particular genetic code are seemingly negatively correlated, make the standard genetic code nearly optimal for the three criteria tested in this work.

Evolutionary Biology

speciesgeocodeR: An R package for linking species occurrences, user-defined regions and phylogenetic trees for biogeography, ecology and evolution

1. Large-scale species occurrence data from geo-referenced observations and collected specimens are crucial for analyses in ecology, evolution and biogeography. Despite the rapidly growing availability of such data, their use in evolutionary analyses is often hampered by tedious manual classification of point occurrences into operational areas, leading to a lack of reproducibility and concerns regarding data quality.\n\n2. Here we present speciesgeocodeR, a user-friendly R-package for data cleaning, data exploration and data visualization of species point occurrences using discrete operational areas, and linking them to analyses invoking phylogenetic trees.\n\n3. The three core functions of the package are 1) automated and reproducible data cleaning, 2) rapid and reproducible classification of point occurrences into discrete operational areas in an adequate format for subsequent biogeographic analyses, and 3) a comprehensive summary and visualization of species distributions to explore large datasets and ensure data quality. In addition, speciesgeocodeR facilitates the access and analysis of publicly available species occurrence data, widely used operational areas and elevation ranges. Other functionalities include the implementation of minimum occurrence thresholds and the visualization of coexistence patterns and range sizes. SpeciesgeocodeR accompanies a richly illustrated and easy-to-follow tutorial and help functions.

Evolutionary Biology

A comparison of one-rate and two-rate inference frameworks for site-specific dN/dS estimation

Two broad paradigms exist for inferring dN/dS, the ratio of nonsynonymous to synonymous substitution rates, from coding sequences: i) a one-rate approach, where dN/dS is represented with a single parameter, or ii) a two-rate approach, where dN and dS are estimated separately. The performances of these two approaches have been well-studied in the specific context of proper model specification, i.e. when the inference model matches the simulation model. By contrast, the relative performances of one-rate vs. two-rate parameterizations when applied to data generated according to a different mechanism remains unclear. Here, we compare the relative merits of one-rate and two-rate approaches in the specific context of model misspecification by simulating alignments with mutation-selection models rather than with dN/dS-based models. We find that one-rate frameworks generally infer more accurate dN/dS point estimates, even when dS varies among sites. In other words, modeling dS variation may substantially reduce accuracy of dN/dS point estimates. These results appear to depend on the selective constraint operating at a given site. In particular, for sites under strong purifying selection (dN/dS<~0.3), one-rate and two-rate models show comparable performances. However, one-rate models significantly outperform two-rate models for sites under moderate-to-weak purifying selection. We attribute this distinction to the fact that, for these more quickly evolving sites, a given substitution is more likely to be nonsynonymous than synonymous. The data will therefore be relatively enriched for nonsynonymous changes, and modeling dS contributes excessive noise to dN/dS estimates. We additionally find that high levels of divergence among sequences, rather than the number of sequences in the alignment, are more critical for obtaining precise point estimates.

Evolutionary Biology

The Implications of Small Stem Cell Niche Sizes and the Distribution of Fitness Effects of New Mutations in Aging and Tumorigenesis

Somatic tissue evolves over a vertebrates lifetime due to the accumulation of mutations in stem cell populations. Mutations may alter cellular fitness and contribute to tumorigenesis or aging. The distribution of mutational effects within somatic cells is not known. Given the unique regulatory regime of somatic cell division we hypothesize that mutational effects in somatic tissue fall into a different framework than whole organisms; one in which there are more mutations of large effect. Through simulation analysis we investigate the fit of tumor incidence curves generated using exponential and power law Distributions of Fitness Effects (DFE) to known tumorigenesis incidence. Modeling considerations include the architecture of stem cell populations, i.e., a large number of very small populations, and mutations that do and do not fix neutrally in the stem cell niche. We find that the typically quantified DFE in whole organisms is sufficient to explain tumorigenesis incidence. Further, due to the effects of small stem cell population sizes, i.e., strong genetic drift, deleterious mutations are predicted to accumulate, resulting in reduced tissue maintenance. Thus, despite there being a large number of stem cells throughout the intestine, its compartmental architecture leads to significant aging, a prime example of Mullers Ratchet.

Evolutionary Biology

Urbanization shapes the demographic history of a native rodent (the white-footed mouse, Peromyscus leucopus) in New York City

How urbanization shapes population genomic diversity and evolution of urban wildlife is largely unexplored. We investigated the impact of urbanization on white-footed mice, Peromyscus leucopus, in the New York City metropolitan area using coalescent-based simulations to infer demographic history from the site frequency spectrum. We assigned individuals to evolutionary clusters and then inferred recent divergence times, population size changes, and migration using genome-wide SNPs genotyped in 23 populations sampled along an urban-to-rural gradient. Both prehistoric climatic events and recent urbanization impacted these populations. Our modeling indicates that post-glacial sea level rise led to isolation of mainland and Long Island populations. These models also indicate that several urban parks represent recently-isolated P. leucopus populations, and the estimated divergence times for these populations are consistent with the history of urbanization in New York City.

Evolutionary Biology

Adaptive Protein Evolution in Animals and the Effective Population Size Hypothesis.

The rate at which genomes adapt to environmental changes and the prevalence of adaptive processes in molecular evolution are two controversial issues in current evolutionary genetics. Previous attempts to quantify the genome-wide rate of adaptation through amino-acid substitution have revealed a surprising diversity of patterns, with some species (e.g. Drosophila) experiencing a very high adaptive rate, while other (e.g. humans) are dominated by nearly-neutral processes. It has been suggested that this discrepancy reflects between-species differences in effective population size. Published studies, however, were mainly focused on model organisms, and relied on disparate data sets and methodologies, so that an overview of the prevalence of adaptive protein evolution in nature is currently lacking. Here we extend existing estimators of the amino-acid adaptive rate by explicitly modelling the effect of favourable mutations on non-synonymous polymorphism patterns, and we apply these methods to a newly-built, homogeneous data set of 44 non-model animal species pairs. Data analysis uncovers a major contribution of adaptive evolution to the amino-acid substitution process across all major metazoan phyla - with the notable exception of humans and primates. The proportion of adaptive amino-acid substitution is found to be positively correlated to species effective population size. This relationship, however, appears to be primarily driven by a decreased rate of nearly-neutral amino-acid substitution due to more efficient purifying selection in large populations. Our results reveal that adaptive processes dominate the evolution of proteins in most animal species, but do not corroborate the hypothesis that adaptive substitutions accumulate at a faster rate in large populations. Implications regarding the factors influencing the rate of adaptive evolution and positive selection detection in humans vs. other organisms are discussed.\n\nAuthor summaryThe rate at which species adapt to environmental changes is a controversial topic. The theory predicts that adaptation is easier in large than in small populations, and the genomic studies of model organisms have revealed a much higher adaptive rate in large population-sized flies than in small population-sized humans and apes. Here we build and analyse a large data set of protein-coding sequences made of thousands of genes in 44 pairs of species from various groups of animals including insects, molluscs, annelids, echinoderms, reptiles, birds, and mammals. Extending and improving existing data analysis methods, we show that adaptation is a major process in protein evolution across all phyla of animals: the proportion of amino-acid substitutions that occurred adaptively is above 50% in a majority of species, and reaches up to 90%. Our analysis does not confirm that population size, here approached through species genetic diversity and ecological traits, does influence the rate of adaptive molecular evolution, but points to human and apes as a special case, compared to other animals, in terms of adaptive genomic processes.

Evolutionary Biology

Interactions retain the co-phylogenetic matching that communities lost

Both species and their interactions are affected by changes that occur at evolutionary time-scales, and these changes shape both ecological communities and their phylogenetic structure. That said, extant ecological community structure is contingent upon random chance, environmental filters, and local effects. It is therefore unclear how much ecological signal local communities should retain. Here we show that, in a host-parasite system where species interactions vary substantially over a continental gradient, the ecological significance of individual interactions is maintained across different scales. Notably, this occurs despite the fact that observed community variation at the local scale frequently tends to weaken or remove community-wide phylogenetic signal. When considered in terms of the interplay between community ecology and coevolutionary theory, our results demonstrate that individual interactions are capable and indeed likely to show a consistent signature of past evolutionary history even when woven into communities that do not.

Evolutionary Biology

Efficient coalescent simulation and genealogical analysis for large sample sizes

A central challenge in the analysis of genetic variation is to provide realistic genome simulation across millions of samples. Present day coalescent simulations do not scale well, or use approximations that fail to capture important long-range linkage properties. Analysing the results of simulations also presents a substantial challenge, as current methods to store genealogies consume a great deal of space, are slow to parse and do not take advantage of shared structure in correlated trees. We solve these problems by introducing sparse trees and coalescence records as the key units of genealogical analysis. Using these tools, exact simulation of the coalescent with recombination for chromosome-sized regions over hundreds of thousands of samples is possible, and substantially faster than present-day approximate methods. We can also analyse the results orders of magnitude more quickly than with existing methods.

Evolutionary Biology

Genome-wide and single-base resolution DNA methylomes of the Sea Lamprey (Petromyzon marinus) Reveal Gradual Transition of the Genomic Methylation Pattern in Early Vertebrates

In eukaryotes, cytosine methylation is a primary heritable epigenetic modification of the genome that regulates many cellular processes. While the whole-genome methylation pattern has been generally conserved in different eukaryotic groups, invertebrates and vertebrates exhibit two distinct patterns. Whereas almost all CpG sites are methylated in most vertebrates, with the exception of short unmethylated regions call CpG islands, the most frequent pattern in invertebrate animals is mosaic methylation, comprising domains of heavily methylated DNA interspersed with domains that are methylation free. The mechanism by which the genome methylation pattern transited from a mosaic to a global pattern and the role of the one or two-round whole-genome duplication in this transition remain largely elusive, partly owing to the lack of methylome data from early vertebrates. In this study, we used the whole-genome bisulfite-sequencing technology to investigate the genome-wide methylation in three tissues (heart, muscle, and sperm) from the sea lamprey, an extant Agarthan vertebrate. Analyses of methylation level and the extent of CpG dinucleotide depletion of geneencoding, intergenic and promoter regions revealed a gradual increase in the methylation level from invertebrates to vertebrates, with the sea lamprey exhibiting an intermediate position. In addition, the methylation level of the majority of CpGs was intermediate in each sea lamprey tissue, indicating a high level of heterogeneity of methylation status between individual cells. In this regard, we defined the genomic methylation pattern of sea lamprey as \"global genomic DNA intermediate methylation\". The methylation features in different genomic regions, such as the transcription start site (TSS) region of the gene body, exon-intron boundaries, transposons, as well as genes grouping with different expression levels, supported the gradual methylation transition hypothesis. We further discussed that the copy number difference in DNA methylation transferases and the loss of the PWWP domain and/or DNTase domain in DNMT3 sub-family enzymes may have contributed to the methylation pattern transition in early vertebrates. These findings demonstrate an intermediate genomic methylation pattern between invertebrates and jawed vertebrates, providing evidence that supports the hypothesis that methylation patterns underwent a gradual transition from invertebrates (mosaic) to vertebrates (global).

Evolutionary Biology

Male sex pheromone components in the butterfly Heliconius melpomene

Sex specific pheromones are known to play an important role in butterfly courtship, and may influence both individual reproductive success and reproductive isolation between species. Extensive ecological, behavioural and genetic studies of Heliconius butterflies have made a substantial contribution to our understanding of speciation. Male pheromones, although long suspected to play an important role, have received relatively little attention in this genus. Here, we combine morphological, chemical, and behavioural analyses of male pheromones in the Neotropical butterfly Heliconius melpomene. First, we identify putative androconia that are specialized brush-like scales that lie within the shiny grey region of the male hindwing. We then describe putative male sex pheromone compounds, which are largely confined to the androconial region of the hindwing of mature males, but are absent in immature males and females. Finally, behavioural choice experiments reveal that females of H. melpomene, H. erato and H. timareta strongly discriminate against conspecific males which have their androconial region experimentally blocked. As well as demonstrating the importance of chemical signalling for female mate choice in Heliconius butterflies, the results describe structures involved in release of the pheromone and a list of potential male sex pheromone compounds.

Evolutionary Biology

Empirical evidence for heterozygote advantage in adapting diploid populations of Saccharomyces cerevisiae

Adaptation in diploids is predicted to proceed via mutations that are at least partially dominant in fitness. Recently we argued that many adaptive mutations might also be commonly overdominant in fitness. Natural (directional) selection acting on overdominant mutations should drive them into the population but then, instead of bringing them to fixation, should maintain them as balanced polymorphisms via heterozygote advantage. If true, this would make adaptive evolution in sexual diploids differ drastically from that of haploids. Unfortunately, the validity of this prediction has not yet been tested experimentally. Here we performed 4 replicate evolutionary experiments with diploid yeast populations (Saccharomyces cerevisiae) growing in glucose-limited continuous cultures. We sequenced 24 evolved clones and identified initial adaptive mutations in all four chemostats. The first adaptive mutations in all four chemostats were three CNVs, all of which proved to be overdominant in fitness. The fact that fitness overdominant mutations were always the first step in independent adaptive walks strongly supports the prediction that heterozygote advantage can arise as a common outcome of directional selection in diploids and demonstrates that overdominance of de novo adaptive mutations in diploids is not rare.

Evolutionary Biology

Accelerated DNA evolution in rats is driven by differential methylation in sperm

Lamarckian inheritance has been largely discredited until the recent discovery of transgenerational epigenetic inheritance. However, transgenerational epigenetic inheritance is still under debate for unable to rule out DNA sequence changes as the underlying cause for heritability. Here, through profiling of the sperm methylomes and genomes of two recently diverged rat subspecies, we analyzed the relationship between epigenetic variation and DNA variation, and their relative contribution to evolution of species. We found that only epigenetic markers located in differentially methylated regions (DMRs) between subspecies, but not within subspecies, can be stably and effectively passed through generations. DMRs in response to both random and stable environmental difference show increased nucleotide diversity, and we demonstrated that it is variance of methylation level but not deamination caused by methylation driving increasing of nucleotide diversity in DMRs, indicating strong relationship between environment-associated changes of chromatin accessibility and increased nucleotide diversity. Further, we detected that accelerated fixation of DNA variants occur only in inter-subspecies DMRs in response to stable environmental difference but not intra-subspecies DMRs in response to random environmental difference or non-DMRs, indicating that this process is possibly driven by environment-associated fixation of divergent methylation status. Our results thus establish a bridge between Lamarckian inheritance and Darwinian selection.

Evolutionary Biology

Evolutionary dynamics of a quantitative trait in a finite asexual population.

In finite populations, mutation limitation and genetic drift can hinder evolutionary diversification. We consider the evolution of a quantitative trait in an asexual population whose size can vary and depends explicitly on the trait. Previous work showed that evolutionary branching is certain (\"deterministic branching\") above a threshold population size, but uncertain (\"stochastic branching\") below it. Using the stationary distribution of the populations trait variance, we identify three qualitatively different sub-domains of \"stochastic branching\" and illustrate our results using a model of social evolution. We find that in very small populations, branching will almost never be observed; in intermediate populations, branching is intermittent, arising and disappearing over time; in larger populations, finally, branching is expected to occur and persist for substantial periods of time. Our study provides a clearer picture of the ecological conditions that facilitate the appearance and persistence of novel evolutionary lineages in the face of genetic drift.

Evolutionary Biology