Search bioRxivSearch

Biology subjects

Sylvain Glemin

Publications and source records attributed to Sylvain Glemin.

4 recordsLinked to original sources

Inference of distribution of fitness effects and proportion of adaptive substitutions from polymorphism data

The distribution of fitness effects (DFE) encompasses deleterious, neutral and beneficial mutations. It conditions the evolutionary trajectory of populations, as well as the rate of adaptive molecular evolution (). Inference of DFE and from patterns of polymorphism (SFS) and divergence data has been a longstanding goal of evolutionary genetics. A widespread assumption shared by numerous methods developed so far to infer DFE and from such data is that beneficial mutations contribute only negligibly to the polymorphism data. Hence, a DFE comprising only deleterious mutations tends to be estimated from SFS data, and is only predicted by contrasting the SFS with divergence data from an outgroup. Here, we develop a hierarchical probabilistic framework that extends on previous methods and also can infer DFE and from polymorphism data alone. We use extensive simulations to examine the performance of our method. We show that both a full DFE, comprising both deleterious and beneficial mutations, and can be inferred without resorting to divergence data. We demonstrate that inference of DFE from polymorphism data alone can in fact provide more reliable estimates, as it does not rely on strong assumptions about a shared DFE between the outgroup and ingroup species used to obtain the SFS and divergence data. We also show that not accounting for the contribution of beneficial mutations to polymorphism data leads to substantially biased estimates of the DFE and . We illustrate these points using our newly developed framework, while also comparing to one of the most widely used inference methods available.

Evolutionary Biology

Limits to adaptation in partially selfing species

In outcrossing populations, \"Haldanes Sieve\" states that recessive beneficial alleles are less likely to fix than dominant ones, because they are less expose to selection when rare. In contrast, selfing organisms are not subject to Haldanes Sieve and are more likely to fix recessive types than outcrossers, as selfing rapidly creates homozygotes, increasing overall selection acting on mutations. However, longer homozygous tracts in selfers also reduces the ability of recombination to create new genotypes. It is unclear how these two effects influence overall adaptation rates in partially selfing organisms. Here, we calculate the fixation probability of beneficial alleles if there is an existing selective sweep in the population. We consider both the potential loss of the second beneficial mutation if it has a weaker advantage than the first, and the possible replacement of the initial allele if the second mutant is fitter. Overall, loss of weaker adaptive alleles during a first selective sweep has a larger impact on preventing fixation of both mutations in highly selfing organisms. Furthermore, the presence of linked mutations has two opposing effects on Haldanes Sieve. First, recessive mutants are disproportionally likely to be lost in outcrossers, so it is likelier that dominant mutations will fix. Second, with elevated rates of adaptive mutation, selective interference annuls the advantage in selfing organisms of not suffering from Haldanes Sieve; outcrossing organisms are more able to fix weak beneficial mutations of any dominance value. Overall, weakened recombination effects can greatly limit adaptation in selfing organisms.

Evolutionary Biology

Introns structure patterns of variation in nucleotide composition in Arabidopsis thaliana and rice protein-coding genes

Plant genomes are large, intron-rich and present a wide range of variation in coding region G + C content. Concerning coding regions, a sort of syndrome can be described in plants: the increase in G + C content is associated with both the increase in heterogeneity among genes within a genome and the increase in variation across genes. Taking advantage of the large number of genes composing plant genomes and the wide range of variation in gene intron number, we performed a comprehensive survey of the patterns of variation in G + C content at different scales from the nucleotide level to the genome scale in two species Arabidopsis thaliana and Oryza sativa, comparing the patterns in genes with different intron numbers. In both species, we observed a pervasive effect of gene intron number and location along genes on G + C content, codon and amino acid frequencies suggesting that in both species, introns have a barrier effect structuring G + C content along genes. In external gene regions (located upstream first or downstream last intron), species-specific factors are shaping G + C content while in internal gene regions (surrounded by introns), G + C content is constrained to remain within a range common to both species. In rice, introns appear as a major determinant of gene G + C content while in A. thaliana introns have a weaker but significant effect. The structuring effect of introns in both species is susceptible to explain the G + C content syndrome observed in plants.

Evolutionary Biology

Quantification of GC-biased gene conversion in the human genome

Many lines of evidence indicate GC-biased gene conversion (gBGC) has a major impact on the evolution of mammalian genomes. However, up to now, this process had not been properly quantified. In principle, the strength of gBGC can be measured from the analysis of derived allele frequency spectra. However, this approach is sensitive to a number of confounding factors. In particular, we show by simulations that the inference is pervasively affected by polymorphism polarization errors, especially at hypermutable sites, and spatial heterogeneity in gBGC strength. Here we propose a new method to quantify gBGC from DAF spectra, incorporating polarization errors and taking spatial heterogeneity into account. This method is very general in that it does not require any prior knowledge about the source of polarization errors and also provides information about mutation patterns. We apply this approach to human polymorphism data from the 1000 genomes project. We show that the strength of gBGC does not differ between hypermutable CpG sites and non-CpG sites, suggesting that in humans gBGC is not caused by the base-excision repair machinery. We further find that the impact of gBGC is concentrated primarily within recombination hotspots: genome-wide, the strength of gBGC is in the nearly neutral area, but 2% of the human genome is subject to strong gBGC, with population-scaled gBGC coefficients above 5. Given that the location of recombination hotspots evolves very rapidly, our analysis predicts that in the long term, a large fraction of the genome is affected by short episodes of strong gBGC.

Evolutionary Biology