Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “evolutionary biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 829 records · Page 46Linked to original sources

Isolation-by-Drift: Quantifying the Respective Contributions of Genetic Drift and Gene Flow in Shaping Spatial Patterns of Genetic Differentiation

O_LIPairwise measures of neutral genetic differentiation are supposed to contain information about past and on-going dispersal events and are thus often used as dependent variables in correlative analyses to elucidate how neutral genetic variation is affected by landscape connectivity. However, spatial heterogeneity in the intensity of genetic drift, stemming from variations in population sizes, may inflate variance in measures of genetic differentiation and lead to erroneous or incomplete interpretations in terms of connectivity. Here, we tested the efficiency of two distance-based metrics designed to capture the unique influence of spatial heterogeneity in local drift on genetic differentiation. These metrics are easily computed from estimates of effective population sizes or from environmental proxies for local carrying capacities, and allow us to introduce the hypothesis of Spatial-Heterogeneity-in-Effective-Population-Sizes (SHNe). SHNe can be tested in a way similar to isolation-by-distance or isolation-by-resistance within the classical landscape genetics hypothesis-testing framework.\nC_LIO_LIWe used simulations under various models of population structure to investigate the reliability of these metrics to quantify the unique contribution of SHNe in explaining patterns of genetic differentiation. We then applied these metrics to an empirical genetic dataset obtained for a freshwater fish (Gobio occitaniae).\nC_LIO_LISimulations showed that SHNe explained up to 60% of variance in genetic differentiation (measured as Fst) in the absence of gene flow, and up to 20% when migration rates were as high as 0.10. Furthermore, one of the two metrics was particularly robust to uncertainty in the estimation of effective population sizes (or proxies for carrying capacity). In the empirical dataset, the effect of SHNe on spatial patterns of Fst was five times higher than that of isolation-by-distance, uniquely contributing to 41% of variance in pairwise Fst. Taking the influence of SHNe into account also allowed decreasing the signal-to-noise ratio, and improving the upper estimate of effective dispersal distance.\nC_LIO_LIWe conclude that the use of SHNe metrics in landscape genetics will substantially improve the understanding of evolutionary drivers of genetic variation, providing substantial information as to the actual drivers of patterns of genetic differentiation in addition to traditional measures of Euclidean distance or landscape resistance.\nC_LI

Evolutionary Biology

Recent demography drives changes in linked selection across the maize genome

Genetic diversity is shaped by the interaction of drift and selection, but the details of this interaction are not well understood. The impact of genetic drift in a population is largely determined by its demographic history, typically summarized by its long-term effective population size (Ne). Rapidly changing population demographics complicate this relationship, however. To better understand how changing demography impacts selection, we used whole-genome sequencing data to investigate patterns of linked selection in domesticated and wild maize (teosinte). We produce the first whole-genome estimate of the demography of maize domestication, showing that maize was reduced to approximately 5% the population size of teosinte before it experienced rapid expansion post-domestication to population sizes much larger than its ancestor. Evaluation of patterns of nucleotide diversity in and near genes shows little evidence of selection on beneficial amino acid substitutions, and that the domestication bottleneck led to a decline in the efficiency of purifying selection in maize. Young alleles, however, show evidence of much stronger purifying selection in maize, reflecting the much larger effective size of present day populations. Our results demonstrate that recent demographic change -- a hallmark of many species including both humans and crops -- can have immediate and wide-ranging impacts on diversity that conflict with would-be expectations based on Ne alone.

Evolutionary Biology

Hybrid dysgenesis in Drosophila simulans associated with a rapid global invasion of the P-element

In a classic example of the invasion of a species by a selfish genetic element, the P-element was horizontally transferred from a distantly related species into Drosophila melanogaster. Despite causing hybrid dysgenesis, a syndrome of abnormal phenotypes that include sterility, the P-element spread globally in the course of a few decades in D. melanogaster. Until recently, its sister species, including D. simulans, remained P-element free. Here, we find a hybrid dysgenesis-like phenotype in the offspring of crosses between D. simulans strains collected in different years; a survey of 181 strains shows that around 20% of strains induce hybrid dysgenesis. Using genomic and transcriptomic data, we show that this dysgenesis-inducing phenotype is associated with the invasion of the P-element. To characterize this invasion temporally and geographically, we survey 631 D. simulans strains collected on three continents and over 27 years for the presence of the P-element. We find that the D. simulans P-element invasion occurred rapidly and nearly simultaneously in the regions surveyed, with strains containing P-elements being rare in 2006 and common by 2014. Importantly, as evidenced by their resistance to the hybrid dysgenesis phenotype, strains collected from the latter phase of this invasion have adapted to suppress the worst effects of the P-element.\n\nAuthor SummarySome genes perform necessary organismal functions, others hijack the cellular machinery to replicate themselves, potentially harming the host in the process. These selfish genes can spread through genomes and species; as a result, eukaryotic genomes are typically saddled with large amounts of parasitic DNA. Here, we chronicle the surprisingly rapid global spread of a selfish transposable element through a close relative of the genetic model, Drosophila melanogaster. We see that, as it spreads, the transposable element is associated with damaging effects, including sterility, but that the flies quickly adapt to the negative consequences of the transposable element.

Evolutionary Biology

Strong Selection is Necessary for Evolution of Blindness in Cave Dwellers

Blindness has evolved repeatedly in cave-dwelling organisms, and investigating the loss of sight in cave dwellers presents an opportunity to understand the operation of fundamental evolutionary processes, including drift, selection, mutation, and migration. The observation of blind organisms has prompted many hypotheses for their blindness, including both accumulation of neutral, loss-of-function mutations and adaptation to darkness. Here we model the evolution of blindness in caves. This model captures the interaction of three forces: (1) selection favoring alleles causing blindness, (2) immigration of sightedness alleles from a surface population, and (3) loss-of-function mutations creating blindness alleles. We investigated the dynamics of this model and determined selection-strength thresholds that result in blindness evolving in caves despite immigration of sightedness alleles from the surface. Our results indicate that strong selection is required for the evolution of blindness in cave-dwelling organisms, which is consistent with recent work suggesting a high metabolic cost of eye development.

Evolutionary Biology

The Genetic Equidistance Phenomenon at the Proteomic Level

The field of molecular evolution started with the alignment of a few protein sequences in the early 1960s. Among the first results found, the genetic equidistance result has turned out to be the most unexpected. It directly inspired the ad hoc universal molecular clock hypothesis that in turn inspired the neutral theory. Unfortunately, however, what is only a maximum distance phenomenon was mistakenly transformed into a mutation rate phenomenon and became known as such. Previous work studied a small set of selected proteins. We have performed proteome wide studies of 7 different sets of proteomes involving a total of 15 species. All 7 sets showed that within each set of 3 species the least complex species is approximately equidistant in average proteome wide identity to the two more complex ones. Thus, the genetic equidis-tance result is a universal phenomenon of maximum distance. There is a reality of constant albeit stepwise or discontinuous increase in complexity during evolution, the rate of which is what the original molecular clock hypothesis is really about. These results provide additional lines of evidence for the recently proposed maximum genetic diversity (MGD) hypothesis.\n\nAvailability and implementationThe source code repository is publicly available at https://github.com/Sephiroth1st/EquidistanceScript\n\nContacthuangshi@sklmg.edu.cn\n\nSupplementary informationSupplementary data are available online.

Evolutionary Biology

A statistical approach to genome size evolution: Observations and explanations

Genome size evolution is a fundamental problem in molecular evolution. Statistical analysis of genome sizes brings new insight into the evolution of genome size. Although the variation of genome sizes is complicated, it is indicated that the genome size evolution can be explained more clearly at taxon level than at species level. I find that the genome size distribution for species in a taxon fits log-normal distribution. And I find a relationship between the phylogeny of life and the statistical features of genome size distributions among taxa. I observed different statistical features of genome size distributions between animal taxa and plant taxa. A log-normal stochastic process model is developed to simulate the genome size evolution. The simulation results on the log-normal distributions of genome sizes and their statistical features agree with the observations.

Evolutionary Biology

Population size and the length of the chromosome blocks identical by descent over generations

In all populations, as the time runs, crossovers break apart ancestor haplotypes, forming smaller blocks at each generation. Some blocks, and eventually all of them, become identical by descent because of the genetic drift. We have in this paper developed and benchmarked a theoretical prediction of the mean length of such blocks and used it to study a simple population model assuming panmixia, no selfing and drift as the only evolutionary pressure. Besides, we have on the one hand derived, for any user defined error threshold, the range of the parameters this prediction is reliable for, and on the other hand shown that the mean length remains constant over time in ideally large populations.

Evolutionary Biology

An in silico comparison of reduced-representation and sequence-capture protocols for phylogenomics

In the age of genome-scale DNA sequencing, choice of molecular marker arguably remains an important decision in planning a phylogenetic study. Using published genomes from 23 primate species, we make a standardized comparison of four of the most frequently used protocols in phylogenomics, viz., targeted sequence-enrichment using ultraconserved element and exon-capture probes, and reduced genomic representation using restriction-site-associated DNA sequencing (RADseq and ddRAD-seq). Here we present a procedure to perform in silico extractions from genomes and create directly comparable datasets for each class of marker. We then compare these datasets in terms of both phylogenetic resolution and ability to consistently and precisely estimate clade ages using fossil-calibrated molecular-clock models. Furthermore, we were also able to directly compare these results to previously published datasets from Sanger-sequenced nuclear exons and mitochondrial genomes under the same analytical conditions. Our results show--although with the exception of the mitochondrial genome and ddRADseq datasets--that for uncontroversial nodes all data classes performed equally well, i.e. they recovered the same well supported topology. However, for one difficult-to-resolve node comprising a rapid diversification (subfamilial relationships among the Cebidae), we report well supported but conflicting topologies among the marker classes, likely the result of mismodelling of gene tree heterogeneity. Likewise, clade age estimates showed consistent discrepancies between datasets; for recent nodes, clade ages estimated by nuclear exon datasets were younger than those of the UCE, RAD and mitochondrial data, but vice versa for the deepest nodes in the primate phylogeny. This effect can be explained by temporal differences in phylogenetic informativeness and choice of clock model used. Finally, we conclude by emphasizing that while huge numbers of loci are probably not required for uncontroversial phylogenetic questions--for which practical considerations such as cost and ease of data generation/sharing/aggregating therefore become increasingly important--accurately modelling heterogeneous data remains as relevant as ever for the more recalcitrant problems.

Evolutionary Biology

Virility does not Imply Immensity: Testis size, Accessory Gland Size and Ejaculate depletion pattern do not Evolve in Response to Experimental Manipulation of Sex Ratio in Drosophila melanogaster

Introduction Introduction Materials and Methods Dissections and Measurements Results Discussion References Females in many species mate more than once and store sperm from more than one male. This leads to post-copulatory competition where sperm from different males compete to fertilize the limited number of eggs produced by the female-typically called Sperm competition (Parker, 1970a) (Wedell et al., 2002). Sperm competition and the resulting post-copulatory selection can significantly alter male reproductive behavior (Simmons et al., 1993) (Cook & Wedell, 1996) (Gage and Barnard, 1996) (Wedell & Cook 1999a,b) (Bretman et al., 2009) (Bretman et al., 2010) ( ...

Evolutionary Biology

Rapid evolution of the inter-sexual genetic correlation for fitness in Drosophila melanogaster

Sexual antagonism (SA) arises when male and female phenotypes are under opposing selection, yet genetically correlated. Until resolved, antagonism limits evolution towards optimal sex-specific phenotypes. Despite its importance for sex-specific adaptation and existing theory, the dynamics of SA resolution are not well understood empirically. Here, we present data from Drosophila melanogaster, compatible with a resolution of SA. We compared two independent replicates of the LHM population in which SA had previously been described. Both had been maintained under identical, controlled conditions, and separated for <250 generations. Although heritabilities of male and female fitness were similar, the inter-sexual genetic correlation differed significantly, being negative in one replicate (indicating SA) but close to zero in the other. Using population sequencing, we show that phenotypic differences were associated with population divergence in allele frequencies at non-random loci across the genome. Large frequency changes were more prevalent in the population without SA and were enriched at loci mapping to genes previously shown to have a sexually antagonistic relationships between expression and fitness. Our data suggest that rapid evolution towards SA resolution has occurred in one of the populations and open avenues towards studying the genetics of SA and its resolution.

Evolutionary Biology

Analysis of the optimality of the Standard Genetic Code.

Many theories have been proposed attempting to explain the origin of the genetic code. While strong reasons remain to believe that the genetic code evolved as a frozen accident, at least for the first few amino acids, other theories remain viable. In this work, we test the optimality of the standard genetic code against approximately 17 million genetic codes, and locate 18 which outperform the standard genetic code at the following three criteria: (a) robustness to point mutation; (b) robustness to frameshift mutation; and (c) ability to encode additional information in the coding region. We use a genetic algorithm to generate and score codes from different parts of the associated landscape, and are, as a result presumably more representative of the entire landscape. Our results show that while the genetic code is sub-optimal for robustness to frameshift mutation and the ability to encode additional information in the coding region, it is very strongly selected for robustness to point mutation. This coupled with the observation that the different performance indicator scores for a particular genetic code are seemingly negatively correlated, make the standard genetic code nearly optimal for the three criteria tested in this work.

Evolutionary Biology

speciesgeocodeR: An R package for linking species occurrences, user-defined regions and phylogenetic trees for biogeography, ecology and evolution

1. Large-scale species occurrence data from geo-referenced observations and collected specimens are crucial for analyses in ecology, evolution and biogeography. Despite the rapidly growing availability of such data, their use in evolutionary analyses is often hampered by tedious manual classification of point occurrences into operational areas, leading to a lack of reproducibility and concerns regarding data quality.\n\n2. Here we present speciesgeocodeR, a user-friendly R-package for data cleaning, data exploration and data visualization of species point occurrences using discrete operational areas, and linking them to analyses invoking phylogenetic trees.\n\n3. The three core functions of the package are 1) automated and reproducible data cleaning, 2) rapid and reproducible classification of point occurrences into discrete operational areas in an adequate format for subsequent biogeographic analyses, and 3) a comprehensive summary and visualization of species distributions to explore large datasets and ensure data quality. In addition, speciesgeocodeR facilitates the access and analysis of publicly available species occurrence data, widely used operational areas and elevation ranges. Other functionalities include the implementation of minimum occurrence thresholds and the visualization of coexistence patterns and range sizes. SpeciesgeocodeR accompanies a richly illustrated and easy-to-follow tutorial and help functions.

Evolutionary Biology

A comparison of one-rate and two-rate inference frameworks for site-specific dN/dS estimation

Two broad paradigms exist for inferring dN/dS, the ratio of nonsynonymous to synonymous substitution rates, from coding sequences: i) a one-rate approach, where dN/dS is represented with a single parameter, or ii) a two-rate approach, where dN and dS are estimated separately. The performances of these two approaches have been well-studied in the specific context of proper model specification, i.e. when the inference model matches the simulation model. By contrast, the relative performances of one-rate vs. two-rate parameterizations when applied to data generated according to a different mechanism remains unclear. Here, we compare the relative merits of one-rate and two-rate approaches in the specific context of model misspecification by simulating alignments with mutation-selection models rather than with dN/dS-based models. We find that one-rate frameworks generally infer more accurate dN/dS point estimates, even when dS varies among sites. In other words, modeling dS variation may substantially reduce accuracy of dN/dS point estimates. These results appear to depend on the selective constraint operating at a given site. In particular, for sites under strong purifying selection (dN/dS<~0.3), one-rate and two-rate models show comparable performances. However, one-rate models significantly outperform two-rate models for sites under moderate-to-weak purifying selection. We attribute this distinction to the fact that, for these more quickly evolving sites, a given substitution is more likely to be nonsynonymous than synonymous. The data will therefore be relatively enriched for nonsynonymous changes, and modeling dS contributes excessive noise to dN/dS estimates. We additionally find that high levels of divergence among sequences, rather than the number of sequences in the alignment, are more critical for obtaining precise point estimates.

Evolutionary Biology

The Implications of Small Stem Cell Niche Sizes and the Distribution of Fitness Effects of New Mutations in Aging and Tumorigenesis

Somatic tissue evolves over a vertebrates lifetime due to the accumulation of mutations in stem cell populations. Mutations may alter cellular fitness and contribute to tumorigenesis or aging. The distribution of mutational effects within somatic cells is not known. Given the unique regulatory regime of somatic cell division we hypothesize that mutational effects in somatic tissue fall into a different framework than whole organisms; one in which there are more mutations of large effect. Through simulation analysis we investigate the fit of tumor incidence curves generated using exponential and power law Distributions of Fitness Effects (DFE) to known tumorigenesis incidence. Modeling considerations include the architecture of stem cell populations, i.e., a large number of very small populations, and mutations that do and do not fix neutrally in the stem cell niche. We find that the typically quantified DFE in whole organisms is sufficient to explain tumorigenesis incidence. Further, due to the effects of small stem cell population sizes, i.e., strong genetic drift, deleterious mutations are predicted to accumulate, resulting in reduced tissue maintenance. Thus, despite there being a large number of stem cells throughout the intestine, its compartmental architecture leads to significant aging, a prime example of Mullers Ratchet.

Evolutionary Biology

Urbanization shapes the demographic history of a native rodent (the white-footed mouse, Peromyscus leucopus) in New York City

How urbanization shapes population genomic diversity and evolution of urban wildlife is largely unexplored. We investigated the impact of urbanization on white-footed mice, Peromyscus leucopus, in the New York City metropolitan area using coalescent-based simulations to infer demographic history from the site frequency spectrum. We assigned individuals to evolutionary clusters and then inferred recent divergence times, population size changes, and migration using genome-wide SNPs genotyped in 23 populations sampled along an urban-to-rural gradient. Both prehistoric climatic events and recent urbanization impacted these populations. Our modeling indicates that post-glacial sea level rise led to isolation of mainland and Long Island populations. These models also indicate that several urban parks represent recently-isolated P. leucopus populations, and the estimated divergence times for these populations are consistent with the history of urbanization in New York City.

Evolutionary Biology

Adaptive Protein Evolution in Animals and the Effective Population Size Hypothesis.

The rate at which genomes adapt to environmental changes and the prevalence of adaptive processes in molecular evolution are two controversial issues in current evolutionary genetics. Previous attempts to quantify the genome-wide rate of adaptation through amino-acid substitution have revealed a surprising diversity of patterns, with some species (e.g. Drosophila) experiencing a very high adaptive rate, while other (e.g. humans) are dominated by nearly-neutral processes. It has been suggested that this discrepancy reflects between-species differences in effective population size. Published studies, however, were mainly focused on model organisms, and relied on disparate data sets and methodologies, so that an overview of the prevalence of adaptive protein evolution in nature is currently lacking. Here we extend existing estimators of the amino-acid adaptive rate by explicitly modelling the effect of favourable mutations on non-synonymous polymorphism patterns, and we apply these methods to a newly-built, homogeneous data set of 44 non-model animal species pairs. Data analysis uncovers a major contribution of adaptive evolution to the amino-acid substitution process across all major metazoan phyla - with the notable exception of humans and primates. The proportion of adaptive amino-acid substitution is found to be positively correlated to species effective population size. This relationship, however, appears to be primarily driven by a decreased rate of nearly-neutral amino-acid substitution due to more efficient purifying selection in large populations. Our results reveal that adaptive processes dominate the evolution of proteins in most animal species, but do not corroborate the hypothesis that adaptive substitutions accumulate at a faster rate in large populations. Implications regarding the factors influencing the rate of adaptive evolution and positive selection detection in humans vs. other organisms are discussed.\n\nAuthor summaryThe rate at which species adapt to environmental changes is a controversial topic. The theory predicts that adaptation is easier in large than in small populations, and the genomic studies of model organisms have revealed a much higher adaptive rate in large population-sized flies than in small population-sized humans and apes. Here we build and analyse a large data set of protein-coding sequences made of thousands of genes in 44 pairs of species from various groups of animals including insects, molluscs, annelids, echinoderms, reptiles, birds, and mammals. Extending and improving existing data analysis methods, we show that adaptation is a major process in protein evolution across all phyla of animals: the proportion of amino-acid substitutions that occurred adaptively is above 50% in a majority of species, and reaches up to 90%. Our analysis does not confirm that population size, here approached through species genetic diversity and ecological traits, does influence the rate of adaptive molecular evolution, but points to human and apes as a special case, compared to other animals, in terms of adaptive genomic processes.

Evolutionary Biology

Interactions retain the co-phylogenetic matching that communities lost

Both species and their interactions are affected by changes that occur at evolutionary time-scales, and these changes shape both ecological communities and their phylogenetic structure. That said, extant ecological community structure is contingent upon random chance, environmental filters, and local effects. It is therefore unclear how much ecological signal local communities should retain. Here we show that, in a host-parasite system where species interactions vary substantially over a continental gradient, the ecological significance of individual interactions is maintained across different scales. Notably, this occurs despite the fact that observed community variation at the local scale frequently tends to weaken or remove community-wide phylogenetic signal. When considered in terms of the interplay between community ecology and coevolutionary theory, our results demonstrate that individual interactions are capable and indeed likely to show a consistent signature of past evolutionary history even when woven into communities that do not.

Evolutionary Biology

Efficient coalescent simulation and genealogical analysis for large sample sizes

A central challenge in the analysis of genetic variation is to provide realistic genome simulation across millions of samples. Present day coalescent simulations do not scale well, or use approximations that fail to capture important long-range linkage properties. Analysing the results of simulations also presents a substantial challenge, as current methods to store genealogies consume a great deal of space, are slow to parse and do not take advantage of shared structure in correlated trees. We solve these problems by introducing sparse trees and coalescence records as the key units of genealogical analysis. Using these tools, exact simulation of the coalescent with recombination for chromosome-sized regions over hundreds of thousands of samples is possible, and substantially faster than present-day approximate methods. We can also analyse the results orders of magnitude more quickly than with existing methods.

Evolutionary Biology