Search bioRxivSearch

Biology subjects

Graham Coop

Publications and source records attributed to Graham Coop.

15 recordsLinked to original sources

Inferring recent demography from isolation by distance of long shared sequence blocks

Recently it has become feasible to detect long blocks of almost identical sequence shared between pairs of genomes. These so called IBD-blocks are direct traces of recent coalescence events, and as such contain ample signal for inferring recent demography. Here, we examine sharing of such blocks in two-dimensional populations with local migration. Using a diffusion approximation to trace genetic ancestry back in time, we derive analytical formulas for patterns of isolation by distance of long IBD-blocks, which can also incorporate recent population density changes. As a main result, we introduce an inference scheme that uses a composite likelihood approach to fit observed block sharing to these formulas. We assess our inference method on simulated block sharing data under several standard population genetics models. We first validate the diffusion approximation by showing that the theoretical results closely match simulated block sharing patterns. We then show that our inference scheme rather accurately and robustly recovers estimates of the dispersal rate and effective density, as well as bounds on recent dynamics of population density. To demonstrate an application, we use our estimation scheme to explore the fit of a diffusion model to Eastern European samples in the POPRES data set. We show that ancestry diffusing with a rate of [Formula] during the last centuries, combined with accelerating population growth, can explain the observed exponential decay of block sharing with pairwise sample distance.

Genetics

Population-genomic inference of the strength and timing of selection against gene flow

The interplay of divergent selection and gene flow is key to understanding how populations adapt to local environments and how new species form. Here, we use DNA polymorphism data and genome-wide variation in recombination rate to jointly infer the strength and timing of selection, as well as the baseline level of gene flow under various demographic scenarios. We model how divergent selection leads to a genome-wide negative correlation between recombination rate and genetic differentiation among populations. Our theory shows that the selection density, i.e. the selection coefficient per base pair, is a key parameter underlying this relationship. We then develop a procedure for parameter estimation that accounts for the confounding effect of background selection. Applying this method to two datasets from Mimulus guttatus, we infer a strong signal of adaptive divergence in the face of gene flow between populations growing on and off phytotoxic serpentine soils. However, the genome-wide intensity of this selection is not exceptional compared to what M. guttatus populations may typically experience when adapting to local conditions. We also find that selection against genome-wide introgression from the selfing sister species M. nasutus has acted to maintain a barrier between these two species over at least the last 250 ky. Our study provides a theoretical framework for linking genome-wide patterns of divergence and recombination with the underlying evolutionary mechanisms that drive this differentiation.

Evolutionary Biology

Estimating time to the common ancestor for a beneficial allele

The haplotypes of a beneficial allele carry information about its history that can shed light on its age and putative cause for its increase in frequency. Specifically, the signature of an alleles age is contained in the pattern of local ancestry that mutation and recombination impose on its haplotypic background. We provide a method to exploit this pattern and infer the time to the common ancestor of a positively selected allele following a rapid increase in frequency. We do so using a hidden Markov model which leverages the length distribution of the shared ancestral haplotype, the accumulation of derived mutations on the ancestral background, and the surrounding background haplotype diversity. Using simulations, we demonstrate how the inclusion of information from both mutation and recombination events increases accuracy relative to approaches that only consider a single type of event. We also show the behavior of the estimator in cases where data do not conform to model assumptions, and provide some diagnostics for assessing and improving inference. Using the method, we analyze population-specific patterns in the 1000 Genomes Project data to provide a global perspective on the timing of adaptation for several variants which show evidence of recent selection and functional relevance to diet, skin pigmentation, and morphology in humans.

Evolutionary Biology

A Genealogical Look at Shared Ancestry on the X Chromosome

Close relatives can share large segments of their genome identical by descent (IBD) that can be identified in genome-wide polymorphism datasets. There are a range of methods to use these IBD segments to identify relatives and estimate their relationship. These methods have focused on sharing on the autosomes, as they provide a rich source of information about genealogical relationships. We can hope to learn additional information about recent ancestry through shared IBD segments on the X chromosome, but currently lack the theoretical framework to use this information fully. Here, we fill this gap by developing probability distributions for the number and length of X chromosome segments shared IBD between an individual and an ancestor k generations back, as well as between half-and full-cousin relationships. Due to the inheritance pattern of the X and the fact that X homologous recombination only occurs in females (outside of the pseudo-autosomal regions), the number of females along a genealogical lineage is a key quantity for understanding the number and length of the IBD segments shared amongst relatives. When inferring relationships among individuals, the number of female ancestors along a genealogical lineage will often be unknown. Therefore, our IBD segment length and number distributions marginalize over this unknown number of recombinational meioses through a distribution of recombinational meioses we derive. We show how our results can be used to estimate the number of female ancestors between two relatives, giving us more genealogical details than possible with autosomal data alone.

Genetics

Does linked selection explain the narrow range of genetic diversity across species?

The relatively narrow range of genetic polymorphism levels across species has been a major source of debate since the inception of molecular population genetics. Recently Corbett-Detig et al found evidence that linked selection strongly constrains levels of polymorphism in species with large census sizes. Here I reexamine this claim and find weak support for this conclusion. While linked selection is an important determinant of polymorphism levels along the genome in many species, we currently lack compelling evidence that it is a major determinant of polymorphism levels among obligately sexual species.

Evolutionary Biology

The Strength of Selection Against Neanderthal Introgression

Hybridization between humans and Neanderthals has resulted in a low level of Neanderthal ancestry scattered across the genomes of many modern-day humans. After hybridization, on average, selection appears to have removed Neanderthal alleles from the human population. Quantifying the strength and causes of this selection against Neanderthal ancestry is key to understanding our relationship to Neanderthals and, more broadly, how populations remain distinct after secondary contact. Here, we develop a novel method for estimating the genome-wide average strength of selection and the density of selected sites using estimates of Neanderthal allele frequency along the genomes of modern-day humans. We confirm that East Asians had somewhat higher initial levels of Neanderthal ancestry than Europeans even after accounting for selection. We find that the bulk of purifying selection against Neanderthal ancestry is best understood as acting on many weakly deleterious alleles. We propose that the majority of these alleles were effectively neutral--and segregating at high frequency--in Neanderthals, but became selected against after entering human populations of much larger effective size. While individually of small effect, these alleles potentially imposed a heavy genetic load on the early-generation human-Neanderthal hybrids. This work suggests that differences in effective population size may play a far more important role in shaping levels of introgression than previously thought.\n\nAuthor SummaryA small percentage of Neanderthal DNA is present in the genomes of many contemporary human populations due to hybridization tens of thousands of years ago. Much of this Neanderthal DNA appears to be deleterious in humans, and natural selection is acting to remove it. One hypothesis is that the underlying alleles were not deleterious in Neanderthals, but rather represent genetic incompatibilities that became deleterious only once they were introduced to the human population. If so, reproductive barriers must have evolved rapidly between Neanderthals and humans after their split. Here, we show that oberved patterns of Neanderthal ancestry in modern humans can be explained simply as a consequence of the difference in effective population size between Neanderthals and humans. Specifically, we find that on average, selection against individual Neanderthal alleles is very weak. This is consistent with the idea that Neanderthals over time accumulated many weakly deleterious alleles that in their small population were effectively neutral. However, after introgressing into larger human populations, those alleles became exposed to purifying selection. Thus, rather than being the result of hybrid incompatibilities, differences between human and Neanderthal effective population sizes appear to have played a key role in shaping our present-day shared ancestry.

Evolutionary Biology

Adaptation to heavy-metal contaminated environments proceeds via selection on pre-existing genetic variation

Anthropogenic environmental changes create evolutionary pressures on populations to adapt to novel stresses. It is as yet unclear, when populations respond to these selective pressures, the extent to which this results in convergent genetic evolution and whether convergence is due to independent mutations or shared ancestral variation. We address these questions using a classic example of adaptation by natural selection by investigating the rapid colonization of the plant species Mimulus guttatus to copper contaminated soils. We use field-based reciprocal transplant experiments to demonstrate that mine alleles at a major copper tolerance locus, Tol1, are strongly selected in the mine environment. We assemble the genome of a mine adapted genotype and identify regions of this genome in tight genetic linkage to Tol1. We discover a set of a multicopper oxidase genes that are genetically linked to Tol1 and exhibit large differences in expression between tolerant and non-tolerant genotypes. We overexpressed this gene in M. guttatus and A. thaliana and found the introduced gene contributes to enhanced copper tolerance. We identify convergent adaptation loci that are additional to Tol1 by measuring genome-wide differences in allele frequency between pairs of mine and off-mine populations and narrow these regions to specific candidate genes using differences in protein sequence and gene expression. Furthermore, patterns of genetic variation at the two most differentiated candidate loci are consistent with selection acting upon alleles that predates the existence of the copper mine habitat. These results suggest that adaptation to the mine habitat occurred via selection on ancestral variation, rather than independent de novo mutations or migration between populations.

Evolutionary Biology

A Coalescent Model of a Sweep from a Uniquely Derived Standing Variant

The use of genetic polymorphism data to understand the dynamics of adaptation and identify the loci that are involved has become a major pursuit of modern evolutionary genetics. In addition to the classical \"hard sweep\" hitchhiking model, recent research has drawn attention to the fact that the dynamics of adaptation can play out in a variety of different ways, and that the specific signatures left behind in population genetic data may depend somewhat strongly on these dynamics. One particular model for which a large number of empirical examples are already known is that in which a single derived mutation arises and drifts to some low frequency before an environmental change causes the allele to become beneficial and sweeps to fixation. Here, we pursue an analytical investigation of this model, bolstered and extended via simulation study. We use coalescent theory to develop an analytical approximation for the effect of a sweep from standing variation on the genealogy at the locus of the selected allele and sites tightly linked to it. We show that the distribution of haplotypes that the selected allele is present on at the time of the environmental change can be approximated by considering recombinant haplotypes as alleles in the infinite alleles model. We show that this approximation can be leveraged to make accurate predictions regarding patterns of genetic polymorphism following such a sweep. We then use simulations to highlight which sources of haplotypic information are likely to be most useful in distinguishing this model from neutrality, as well as from other sweep models, such as the classic hard sweep, and multiple mutation soft sweeps. We find that in general, adaptation from a uniquely derived standing variant will be difficult to detect on the basis of genetic polymorphism data alone, and when it can be detected, it will be difficult to distinguish from other varieties of selective sweeps.

Evolutionary Biology

The Spatial Mixing of Genomes in Secondary Contact Zones

Recent genomic studies have highlighted the important role of admixture in shaping genome-wide patterns of diversity. Past admixture leaves a population genomic signature of linkage disequilibrium (LD), reflecting the mixing of parental chromosomes by segregation and recombination. The extent of this LD can be used to infer the timing of admixture. However, the results of inference can depend strongly on the assumed demographic model. Here, we introduce a theoretical framework for modeling patterns of LD in a geographic contact zone where two differentiated populations are diffusing back together. We derive expressions for the expected LD and admixture tract lengths across geographic space as a function of the age of the contact zone and the dispersal distance of individuals. We develop an approach to infer age of contact zones using population genomic data from multiple spatially sampled populations by fitting our model to the decay of LD with recombination distance. We use our approach to explore the fit of a geographic contact zone model to three human population genomic datasets from populations along the Indonesian archipelago, populations in Central Asia and populations in India.

Evolutionary Biology

A Spatial Framework for Understanding Population Structure and Admixture.

Geographic patterns of genetic variation within modern populations, produced by complex histories of migration, can be difficult to infer and visually summarize. A general consequence of geographically limited dispersal is that samples from nearby locations tend to be more closely related than samples from distant locations, and so genetic covariance often recapitulates geographic proximity. We use genome-wide polymorphism data to build \"geogenetic maps,\" which, when applied to stationary populations, produces a map of the geographic positions of the populations, but with distances distorted to reflect historical rates of gene flow. In the underlying model, allele frequency covariance is a decreasing function of geogenetic distance, and nonlocal gene flow such as admixture can be identified as anomalously strong covariance over long distances. This admixture is explicitly co-estimated and depicted as arrows, from the source of admixture to the recipient, on the geogenetic map. We demonstrate the utility of this method on a circum-Tibetan sampling of the greenish warbler (Phylloscopus trochiloides), in which we find evidence for gene flow between the adjacent, terminal populations of the ring species. We also analyze a global sampling of human populations, for which we largely recover the geography of the sampling, with support for significant histories of admixture in many samples. This new tool for understanding and visualizing patterns of population structure is implemented in a Bayesian framework in the program SpaceMix.\n\nAuthor SummaryIn this paper, we introduce a statistical method for inferring, for a set of sequenced samples, a map in which the distances between population locations reflect genetic, rather than geographic, proximity. Two populations that are sampled at distant locations but that are genetically similar (perhaps one was recently founded by a colonization event from the other) may have inferred locations that are nearby, while two populations that are sampled close together, but that are genetically dissimilar (e.g., are separated by a barrier), may have inferred locations that are farther apart. The result is a \"geogenetic\" map in which the distances between populations are effective distances, indicative of the way that populations perceive the distances between themselves: the \"organisms-eye view\" of the world. Added to this, \"admixture\" can be thought of as the outcome of unusually long-distance gene flow; it results in relatedness between populations that is anomalously high given the distance that separates them. We depict the effect of admixture using arrows, from a source of admixture to its target, on the inferred map. The inferred geogenetic map is an intuitive and information-rich visual summary of patterns of population structure.

Evolutionary Biology

The role of standing variation in geographic convergent adaptation

The extent to which populations experiencing shared selective pressures adapt through a shared genetic response is relevant to many questions in evolutionary biology. In a number of well studied traits and species, it appears that convergent evolution within species is common. In this paper, we explore how standing, genetic variation contributes to convergent genetic responses in a geographically spread population, extending our previous work on the topic. Geographically limited dispersal slows the spread of each selected allele, hence allowing other alleles - newly arisen mutants or present as standing variation - to spread before any one comes to dominate the population. When such alleles meet, their progress is substantially slowed - if the alleles are selectively equivalent, they mix slowly, dividing the species range into a random tessellation, which can be well understood by analogy to a Poisson process model of crystallization. In this framework, we derive the geographic scale over which a typical allele is expected to dominate, the time it takes the species to adapt as a whole, and the proportion of adaptive alleles that arise from standing variation. Finally, we explore how negative pleiotropic effects of alleles before an environment change can bias the subset of alleles that contribute to the species adaptive response. We apply the results to the many geographically localized G6PD deficiency alleles thought to confer resistance to malaria, where the large mutational target size makes it a likely candidate for adaptation from standing variation, despite the selective cost of G6PD deficiency alleles in the absence of malaria. We find the numbers and geographic spread of these alleles matches our predictions reasonably well, consistent with the view that they arose from a combination of standing variation and new mutations since the advent of malaria. Our results suggest that much of adaptation may be geographically local even when selection pressures are homogeneous. Therefore, we argue that caution must be exercised when arguing that strongly geographically restricted alleles are necessarily the outcome of local adaptation. We close by discussing the implications of these results for ideas of species coherence and the nature of divergence between species.

Evolutionary Biology

Convergent Evolution During Local Adaptation to Patchy Landscapes

Species often encounter, and adapt to, many patches of locally similar environmental conditions across their range. Such adaptation can occur through convergent evolution as different alleles arise and spread in different patches, or through the spread of alleles by migration acting to synchronize adaptation across the species. The tension between the two reflects the degree of constraint imposed on evolution by the underlying genetic architecture versus how effectively selection acts to inhibit the geographic spread of locally adapted alleles. This paper studies a model of the balance between these two routes to adaptation in continuous environments with patchy selection pressures. We address the following questions: How long does it take for a novel, locally adapted allele to appear in a patch of habitat where it is favored through mutation? Or, through migration from another, already adapted patch? Which is more likely to occur, as a function of distance between the patches? How can we tell which has occurred, i.e. what population genetic signal is left by the spread of migrant alleles and how long does this signal persist for? To answer these questions we decompose the migration-selection equilibrium surrounding an already adapted patch into families of migrant alleles, in particular treating those rare families that reach new patches as spatial branching processes. This provides a way to understand the role of geographic separation between patches in promoting convergent adaptation and the genomic signals it leaves behind. We illustrate these ideas using the convergent evolution of cryptic coloration in the rock pocket mouse, Chaetodipus intermedius, as an empirical example.\n\nAuthor SummaryOften, a large species range will include patches where the species differs because it has adapted to locally differing conditions. For instance, rock pocket mice are often found with a coat color that matches the rocks they live in, and the difference in coloration is known to be controlled genetically. Sometimes, similar genetic changes have occurred independently in different patches, suggesting that there were few accessible ways to evolve the locally adaptive form. However, the genetic basis could also be shared if rare migrants may carry the locally beneficial genotypes between nearby patches, despite being at a disadvantage between the patches. We use a mathematical model of random migration to determine how quickly adaptation is expected to occur through new mutation or migration from other patches, and study in more detail what we would expect successful migrations between patches to look like. The results are useful for determining whether similar adaptations in different locations are likely to have the same genetic basis or not, and more generally in understanding how species adapt to patchy, heterogeneous landscapes.

Evolutionary Biology

Sperm should evolve to make female meiosis fair.

Genomic conflicts arise when an allele gains an evolutionary advantage at a cost to organismal fitness. Oogenesis is inherently susceptible to such conflicts because alleles compete for inclusion into the egg. Alleles that distort meiosis in their favor (i.e. meiotic drivers) often decrease organismal fitness, and therefore indirectly favor the evolution of mechanisms to suppress meiotic drive. In this light, many facets of oogenesis and gametogenesis have been interpreted as mechanisms of protection against genomic outlaws. That females of many animal species do not complete meiosis until after fertilization, appears to run counter to this interpretation, because this delay provides an opportunity for sperm-acting alleles to meddle with the outcome of female meiosis and help like alleles drive in heterozygous females. Contrary to this perceived danger, the population genetic theory presented herein suggests that, in fact, sperm nearly always evolve to increase the fairness of female meiosis in the face of genomic conflicts. These results are consistent with the apparent sperm dependence of the best characterized female meiotic drivers in animals. Rather than providing an opportunity for sperm collaboration in female meiotic drive, the fertilization requirement indirectly protects females from meiotic drivers by providing sperm an opportunity to suppress drive.

Evolutionary Biology

A Population Genetic Signature of Polygenic Local Adaptation

Adaptation in response to selection on polygenic phenotypes occurs via subtle allele frequencies shifts at many loci. Current population genomic techniques are not well posed to identify such signals. In the past decade, detailed knowledge about the specific loci underlying polygenic traits has begun to emerge from genome-wide association studies (GWAS). Here we combine this knowledge from GWAS with robust population genetic modeling to identify traits that have undergone local adaptation. Using GWAS data, we estimate the mean additive genetic value for a give phenotype across many populations as simple weighted sums of allele frequencies. We model the expected differentiation of GWAS loci among populations under neutrality to develop simple tests of selection across an arbitrary number of populations with arbitrary population structure. To find support for the role of specific environmental variables in local adaptation we test for correlations with the estimated genetic values. We also develop a general test of local adaptation to identify overdispersion of the estimated genetic values values among populations. This test is a natural generalization of QST /FST comparisons based on GWAS predictions. Finally we lay out a framework to identify the individual populations or groups of populations that contribute to the signal of overdispersion. These tests have considerably greater power than their single locus equivalents due to the fact that they look for positive covariance between like effect alleles. We apply our tests to the human genome diversity panel dataset using GWAS data for six different traits. This analysis uncovers a number of putative signals of local adaptation, and we discuss the biological interpretation and caveats of these results.

Genetics

Speciation and introgression between Mimulus nasutus and Mimulus guttatus

Mimulus guttatus and M. nasutus are an evolutionary and ecological model sister species pair differentiated by ecology, mating system, and partial reproductive isolation. Despite extensive research on this system, the history of divergence and differentiation in this sister pair is unclear. We present and analyze a novel population genomic data set which shows that M. nasutus "budded" off of a central Californian M. guttatus population within the last 200 to 500 thousand years. In this time, the M. nasutus genome has accrued numerous genomic signatures of the transition to predominant selfing. Despite clear biological differentiation, we document ongoing, bidirectional introgression. We observe a negative relationship between the recombination rate and divergence between M. nasutus and sympatric M. guttatus samples, suggesting that selection acts against M. nasutus ancestry in M. guttatus.

Evolutionary Biology