Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “evolutionary biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Comparing RADseq and microsatellites to infer complex phylogeographic patterns, a real data informed perspective in the Crucian carp, Carassius carassius, L.

The conservation of threatened species must be underpinned by phylogeographic knowledge in order to be effective. This need is epitomised by the freshwater fish Carassius carassius, which has recently undergone drastic declines across much of its European range. Restriction Site Associated DNA sequencing (RADseq) is being increasingly used for such phylogeographic questions, however RADseq is expensive, and limitations on sample number must be weighed against the benefit of large numbers of markers. Such tradeoffs have predominantly been addressed using simulated data. Here we compare the results generated from microsatellites and RADseq to the phylogeography of C. carassius, to add real-data-informed perspectives to this important debate. These datasets, along with data from the mitochondrial cytochrome b gene, agree on broad phylogeographic patterns; showing the existence of two previously unidentified C. carassius lineages in Europe. These lineages have been isolated for approximately 2.2-2.3 M years, and should arguably be considered as separate conservation units. RADseq recovered finer population structure and stronger patterns of IBD than microsatellites, despite including only 17.6% of samples (38% of populations and 52% of samples per population). RADseq was also used along with Approximate Bayesian Computation to show that the postglacial colonisation routes of C. carassius differ from the general patterns of freshwater fish in Europe, likely as a result of their distinctive ecology.

Evolutionary Biology

Evolutionary dynamics of roX lncRNA function and genomic occupancy

Many long noncoding RNAs (lncRNAs) can regulate chromatin states, but the evolutionary origin and dynamics driving lncRNA-genome interactions are unclear. We developed an integrative strategy that identifies lncRNA orthologs in different species despite limited sequence similarity that is applicable to fly and mammalian lncRNAs. Analysis of the roX lncRNAs, which are essential for dosage compensation of the single X-chromosome in Drosophila males, revealed 47 new roX orthologs in diverse Drosophilid species across ~40 million years of evolution. Genetic rescue by roX orthologs and engineered synthetic lncRNAs showed that evolutionary maintenance of focal structural repeats mediates roX function. Genomic occupancy maps of roX RNAs in four species revealed rapid turnover of individual binding sites but conservation within nearby chromosomal neighborhoods. Many new roX binding sites evolved from DNA encoding a pre-existing RNA splicing signal, effectively linking dosage compensation to transcribed genes. Thus, evolutionary analysis illuminates the principles for the birth and death of lncRNAs and their genomic targets.

Evolutionary Biology

Evolutionary quantitative genomics of Populus trichocarpa

Forest trees generally show high levels of local adaptation and efforts focusing on understanding adaptation to climate will be crucial for species survival and management.\n\nMerging quantitative genetics and population genomics, we studied the molecular basis of climate adaptation in 433 Populus trichocarpa (black cottonwood) genotypes originating across western North America. Variation in 74 field-assessed traits (growth, ecophysiology, phenology, leaf stomata, wood, and disease resistance) was investigated for signatures of selection (comparing QST-FST) using clustering of individuals by climate of origin. 29,354 SNPs were investigated employing three different outlier detection methods.\n\nNarrow-sense QST for 53% of distinct field QST traits was significantly divergent from expectations of neutrality (indicating adaptive trait variation); 2,855 SNPs showed signals of diversifying selection and of these, 118 SNPs (within 81 genes) were associated with adaptive traits (based on significant QST). Many SNPs were putatively pleiotropic for functionally uncorrelated adaptive traits, such as autumn phenology, height, and disease resistance.\n\nEvolutionary quantitative genomics in P. trichocarpa provides an enhanced understanding regarding the molecular basis of climate-driven selection in forest trees. We highlight that important loci underlying adaptive trait variation also show relationship to climate of origin.\n\nAuthor summaryComparisons between population differentiation on the basis of quantitative traits and neutral genetic markers inform about the importance of natural selection, genetic drift and gene flow for local adaptation of populations. Here, we address fundamental questions regarding the molecular basis of adaptation in undomesticated forest tree populations to past climatic environments by employing an integrative quantitative genetics and landscape genomics approach. Marker-inferred relatedness was estimated to obtain the narrow-sense estimate of population differentiation in wild populations. We analyzed an unstructured population of common garden grown Populus trichocarpa individuals to uncover different extents of variation for a suite of field traits, wood quality and pathogen resistance with temperature and precipitation. We consider our approach the most comprehensive, as it uncovers the molecular mechanisms of adaptation using multiple methods and tests. We provide a detailed outline of the required analyses for studying adaptation to the environment in a population genomics context to better understand the species potential adaptive capacity to future climatic scenarios.

Evolutionary Biology

Genetic evidence challenges the native status of a threatened freshwater fish (Carassius carassius) in England

A fundamental consideration for the conservation of a species is the extent of its native range, however defining a native range is often challenging as changing environments drive shifts in species distributions over time. The crucian carp, Carassius carassius (L.) is a threatened freshwater fish native to much of Europe, however the extent of this range is ambiguous. One particularly contentious region is England, in which C. carassius is currently considered native on the basis of anecdotal evidence. Here, we use 13 microsatellite loci, population structure analyses and approximate bayesian computation (ABC), to empirically test the native status of C. carassius in England. Contrary to the current consensus, ABC yields strong support for introduced origins of C. carassius in England, with posterior distribution estimates placing their introduction in the 15th century, well after the loss of the doggerland landbridge. This result brings to light an interesting and timely debate surrounding our motivations for the conservation of species. We discuss this topic, and make arguments for the continued conservation of C. carassius in England, despite its non-native origins.

Evolutionary Biology

Limits to adaptation in partially selfing species

In outcrossing populations, \"Haldanes Sieve\" states that recessive beneficial alleles are less likely to fix than dominant ones, because they are less expose to selection when rare. In contrast, selfing organisms are not subject to Haldanes Sieve and are more likely to fix recessive types than outcrossers, as selfing rapidly creates homozygotes, increasing overall selection acting on mutations. However, longer homozygous tracts in selfers also reduces the ability of recombination to create new genotypes. It is unclear how these two effects influence overall adaptation rates in partially selfing organisms. Here, we calculate the fixation probability of beneficial alleles if there is an existing selective sweep in the population. We consider both the potential loss of the second beneficial mutation if it has a weaker advantage than the first, and the possible replacement of the initial allele if the second mutant is fitter. Overall, loss of weaker adaptive alleles during a first selective sweep has a larger impact on preventing fixation of both mutations in highly selfing organisms. Furthermore, the presence of linked mutations has two opposing effects on Haldanes Sieve. First, recessive mutants are disproportionally likely to be lost in outcrossers, so it is likelier that dominant mutations will fix. Second, with elevated rates of adaptive mutation, selective interference annuls the advantage in selfing organisms of not suffering from Haldanes Sieve; outcrossing organisms are more able to fix weak beneficial mutations of any dominance value. Overall, weakened recombination effects can greatly limit adaptation in selfing organisms.

Evolutionary Biology

Overlapping Genes and Size Constraints in Viruses - An Evolutionary Perspective

Viruses are the simplest replicating units, characterized by a limited number of coding genes and an exceptionally high rate of overlapping genes. We sought a unified explanation for the evolutionary constraints that govern genome sizes, gene overlapping and capsid properties. We performed an unbiased statistical analysis over the [~]100 known viral families, and came to refute widespread assumptions regarding viral evolution. We found that the volume utilization of viral capsids is often low, and greatly varies among families. Most notably, we show that the total amount of gene overlapping is tightly bounded. Although viruses expand three orders of magnitude in genome length, their absolute amount of gene overlapping almost never exceeds 1500 nucleotides, and mostly confined to <4 significant overlapping instances. Our results argue against the common theory by which gene overlapping is driven by a necessity of viruses to compress their genome. Instead, we support the notion that overlapping has a role in gene novelty and evolution exploration.

Evolutionary Biology

Selection for mitochondrial quality drives the evolution of sexes with a dedicated germline

The origin of the germline-soma distinction is a fundamental unsolved question. Plants and basal metazoans do not have a germline but generate gametes from somatic tissues (somatic gametogenesis), whereas most bilaterians sequester a germline. We develop an evolutionary model which shows that selection for mitochondrial quality drives germline evolution. In organisms with low mitochondrial mutation rates, segregation of mutations over multiple cell divisions generates variation, allowing selection to optimize gamete quality through somatic gametogenesis. Higher mutation rates promote early germline sequestration. Oogamy reduces mitochondrial segregation in early development, improving adult fitness by restricting variation between tissues, but also limiting variation between early-sequestered oocytes, undermining gamete quality. Oocyte variation is restored through proliferation and random culling (atresia) of precursor cells. We predict a novel pathway from basal metazoans lacking a germline to active bilaterians with early sequestration of large oocytes subject to atresia, allowing the emergence of complex developmental processes.

Evolutionary Biology

Natural selection and recombination rate variation shape nucleotide polymorphism across the genomes of three related Populus species.

A central aim of evolutionary genomics is to identify the relative roles that various evolutionary forces have played in generating and shaping genetic variation within and among species. Here we use whole-genome re-sequencing data to characterize and compare genome-wide patterns of nucleotide polymorphism, site frequency spectrum and population-scaled recombination rates in three species of Populus: P. tremula, P. tremuloides and P. trichocarpa. We find that P. tremuloides has the highest level of genome-wide variation, skewed allele frequencies and population-scaled recombination rates, whereas P. trichocarpa harbors the lowest. Our findings highlight multiple lines of evidence suggesting that natural selection, both due to purifying and positive selection, has widely shaped patterns of nucleotide polymorphism at linked neutral sites in all three species. Differences in effective population sizes and rates of recombination are largely explaining the disparate magnitudes and signatures of linked selection we observe among species. The present work provides the first phylogenetic comparative study at genome-wide scale in forest trees. This information will also improve our ability to understand how various evolutionary forces have interacted to influence genome evolution among related species.

Evolutionary Biology

HacDivSel: Two new methods (haplotype-based and outlier-based) for the detection of divergent selection in pairs of populations

The detection of genomic regions involved in local adaptation is an important topic in current population genetics. There are several detection strategies available depending on the kind of genetic and demographic information at hand. A common drawback is the high risk of false positives. In this study, we introduce two complementary methods for the detection of divergent selection from populations connected by migration. Both methods have been developed with the aim of being robust to false positives. The first method combines haplotype information with inter-population differentiation (FST). Evidence of divergent selection is concluded only when both the haplotype pattern and the FST value support it. The second method is developed for independently segregating markers i.e. there is no haplotype information at hand. In this case, the power to detect selection is attained by developing a new outlier test based on detecting a bimodal distribution. The test computes the FST outliers and then assumes that those of interest would have a different mode which is detected by a clustering algorithm. The utility of the two methods is demonstrated through simulations and the analysis of real data. The simulation results show power ranging from 60-94% in several of the scenarios whilst the false positive rate is controlled below the nominal level in every scenario. The analysis of real samples consisted of phased data from the HapMap project and unphased data from intertidal marine snail ecotypes. The software HacDivSel implements the methods explained in this manuscript.

Evolutionary Biology

General methods for evolutionary quantitative genetic inference from generalised mixed models.

Methods for inference and interpretation of evolutionary quantitative genetic parameters, and for prediction of the response to selection, are best developed for traits with normal distributions. Many traits of evolutionary interest, including many life history and behavioural traits, have inherently non-normal distributions. The generalised linear mixed model (GLMM) framework has become a widely used tool for estimating quantitative genetic parameters for non-normal traits. However, whereas GLMMs provide inference on a statistically-convenient latent scale, it is sometimes desirable to express quantitative genetic parameters on the scale upon which traits are expressed. The parameters of a fitted GLMMs, despite being on a latent scale, fully determine all quantities of potential interest on the scale on which traits are expressed. We provide expressions for deriving each of such quantities, including population means, phenotypic (co)variances, variance components including additive genetic (co)variances, and parameters such as heritability. We demonstrate that fixed effects have a strong impact on those parameters and show how to deal for this effect by averaging or integrating over fixed effects. The expressions require integration of quantities determined by the link function, over distributions of latent values. In general cases, the required integrals must be solved numerically, but efficient methods are available and we provide an implementation in an R package, QGglmm. We show that known formulae for quantities such as heritability of traits with Binomial and Poisson distributions are special cases of our expressions. Additionally, we show how fitted GLMM can be incorporated into existing methods for predicting evolutionary trajectories. We demonstrate the accuracy of the resulting method for evolutionary prediction by simulation, and apply our approach to data from a wild pedigreed vertebrate population.

Evolutionary Biology

Divergent MLS1 promoters lie on a fitness plateau for gene expression

Qualitative patterns of gene activation and repression are often conserved despite an abundance of quantitative variation in expression levels within and between species. A major challenge to interpreting patterns of expression divergence is knowing which changes in gene expression affect fitness. To characterize the fitness effects of gene expression divergence we placed orthologous promoters from eight yeast species upstream of malate synthase (MLS1) in Saccharomyces cerevisiae. As expected, we found these promoters varied in their expression level under activated and repressed conditions as well as in their dynamic response following loss of glucose repression. Despite these differences, only a single promoter driving near basal levels of expression caused a detectable loss of fitness. We conclude that the MLS1 promoter lies on a fitness plateau whereby even large changes in gene expression can be tolerated without a substantial loss of fitness.

Evolutionary Biology

Selection on heritable heterozygosity but no response to selection. Why?

The realisation that heterozygosity can be heritable has recently generated some elegant research. However, none of this work has discussed the fact that when heterozygote advantage occurs, heterozygosity can be heritable, yet allele frequencies remain at equilibrium and do not evolve with time. From a quantitative genetic perspective this means the character is heritable, is under selection, yet no response to selection is observed. We explain why this is the case, and discuss potential implications for the study of evolution in the wild.

Evolutionary Biology

EvolQG - An R package for evolutionary quantitative genetics

We present an open source package for performing evolutionary quantitative genetics analyses in the R environment for statistical computing. Evolutionary theory shows that evolution depends critically on the available variation in a given population. When dealing with many quantitative traits this variation is expressed in the form of a covariance matrix, particularly the additive genetic covariance matrix or sometimes the phenotypic matrix, when the genetic matrix is unavailable. Given this mathematical representation of available variation, the EvolQG package provides functions for calculation of relevant evolutionary statistics, estimation of sampling error, corrections for this error, matrix comparison via correlations and distances, and functions for testing evolutionary hypotheses on taxa diversification.

Evolutionary Biology

Inference of complex population histories using whole-genome sequences from multiple populations

There has been much interest in analyzing genome-scale DNA sequence data to infer population histories, but inference methods developed hitherto are limited in model complexity and computational scalability. Here we present an efficient, flexible statistical method, diCal2, that can utilize whole-genome sequence data from multiple populations to infer complex demographic models involving population size changes, population splits, admixture, and migration. Applying our method to data from Australian, East Asian, European, and Papuan populations, we find that the population ancestral to Australians and Papuans started separating from East Asians and Europeans about 100,000 years ago, and that the separation of East Asians and Europeans started about 50,000 years ago, with pervasive gene flow between all pairs of populations.

Evolutionary Biology

Mapping phylogenetic trees to reveal distinct patterns of evolution

Evolutionary relationships are described by phylogenetic trees, but a central barrier in many fields is the difficulty of interpreting data containing conflicting phylogenetic signals. Obtaining credible trees that capture the relationships present in complex data is one of the fundamental challenges in evolution today. We present a way to map trees which extracts distinct alternative evolutionary relationships embedded in data and resolves phylogenetic uncertainty. Our method reveals a remarkably distinct phylogenetic signature in the VP30 gene of Ebolavirus, indicating possible recombination with significant implications for vaccine development. Moving to higher organisms, we use our approach to detect alternative histories of the evolution of anole lizards. Our approach has the capacity to resolve key areas of uncertainty in the study of evolution, and to broaden the credibility and appeal of phylogenetic methods.

Evolutionary Biology

How sex-biased dispersal affects conflict over parental investment

Abstract Existing models of parental investment have mainly focused on interactions at the level of the family, and have paid much less attention to the impact of population-level processes. Here we extend classical models of parental care to assess the impact of population structure and limited dispersal. We find that sex-differences in dispersal substantially affect the amount of care provided by each parent, with the more philopatric sex providing the majority of the care to young. This effect is most pronounced in highly viscous populations: in such cases, when classical models would predict stable biparental care, inclusion of a modest sex difference in dispersal leads to uniparental care by the philopatric sex. In addition, mating skew also affects sex-differences in parental investment, with the more numerous sex providing most of the care. However, the effect of mating skew only holds when parents care for their own offspring. When individuals breed communally, we recover the previous finding that the more philopatric sex provides most of the care, even when it is the rare sex. Finally, we show that sex-differences in dispersal can mask the existence of sex-specific costs of care, because the philopatric sex may provide most of the care even in the face of far higher mortality costs relative to the dispersing sex. We conclude that sex-biased dispersal is likely to be an important, yet currently overlooked driver of sex-differences in parental care.

Evolutionary Biology

Genetic structure of the stingless bee Tetragonisca angustula

The stingless bee Tetragonisca angustula Latreille 1811 is distributed from Mexico to Argentina and is one of the most widespread bee species in the Neotropics. However, this wide distribution contrasts with the short distance traveled by females to build new nests. Here we evaluate the genetic structure of several populations of T. angustula using mitochondrial DNA and microsatellites. These markers can help us to detect differences in the migratory behavior of males and females. Our results show that the populations are highly differentiated suggesting that both females and males have low dispersal distance. Therefore, its continental distribution probably consists of several cryptic species.

Evolutionary Biology

Inference of multiple-wave population admixture by modeling decay of linkage disequilibrium with multiple exponential functions

Admixture-introduced linkage disequilibrium (LD) has recently been introduced into the inference of the histories of complex admixtures. However, the influence of ancestral source populations on the LD pattern in admixed populations is not properly taken into consideration by currently available methods, which affects the estimation of several gene flow parameters from empirical data. We first illustrated the dynamic changes of LD in admixed populations and mathematically formulated the LD under a generalized admixture model with finite population size. We next developed a new method, MALDmef, by fitting LD with multiple exponential functions for inferring and dating multiple-wave admixtures. MALDmef takes into account the effects of source populations which substantially affect modeling LD in admixed population, which renders it capable of efficiently detecting and dating multiple-wave admixture events. The performance of MALDmef was evaluated by simulation and it was shown to be more accurate than MALDER, a state-of-the-art method that was recently developed for similar purposes, under various admixture models. We further applied MALDmef to analyzing genome-wide data from the Human Genome Diversity Project (HGDP) and the HapMap Project. Interestingly, we were able to identify more than one admixture events in several populations, which have yet to be reported. For example, two major admixture events were identified in the Xinjiang Uyghur, occurring around 27-30 generations ago and 182-195 generations ago, respectively. In an African population (MKK), three recent major admixtures occurring 13-16, 50-67, and 107-139 generations ago were detected. Our method is a considerable improvement over other current methods and further facilitates the inference of the histories of complex population admixtures.

Evolutionary Biology