Search bioRxivSearch

Biology subjects

Coop, G.

Publications and source records attributed to Coop, G..

8 recordsLinked to original sources

Reconstructing the history of polygenic scores using coalescent trees

1Genome-wide association studies (GWAS) have revealed that many traits are highly polygenic, in that their within-population variance is governed in part by small-effect variants at many genetic loci. Standard population-genetic methods for inferring evolutionary history are ill-suited for polygenic traits--when there are many variants of small effect, signatures of natural selection are spread across the genome and subtle at any one locus. In the last several years, several methods have emerged for detecting the action of natural selection on polygenic scores, sums of genotypes weighted by GWAS effect sizes. However, most existing methods do not reveal the timing or strength of selection. Here, we present a set of methods for estimating the historical time course of a population-mean polygenic score using local coalescent trees at GWAS loci. These time courses are estimated by using coalescent theory to relate the branch lengths of trees to allele-frequency change. The resulting time course can be tested for evidence of natural selection. We present theory and simulations supporting our procedures, as well as estimated time courses of polygenic scores for human height. Because of its grounding in coalescent theory, the framework presented here can be extended to a variety of demographic scenarios, and its usefulness will increase as both GWAS and ancestral recombination graph (ARG) inference continue to progress.

evolutionary biology

Allele frequency dynamics in a pedigreed natural population

A central goal of population genetics is to understand how genetic drift, natural selection, and gene flow shape allele frequencies through time. However, the actual processes underlying these changes - variation in individual survival, reproductive success, and movement - are often difficult to quantify. Fully understanding these processes requires the population pedigree, the set of relationships among all individuals in the population through time. Here, we use extensive pedigree and genomic information from a long-studied natural population of Florida Scrub-Jays (Aphelocoma coerulescens) to directly characterize the relative roles of different evolutionary processes in shaping patterns of genetic variation through time. We performed gene dropping simulations to estimate individual genetic contributions to the population and model drift on the known pedigree. We found that observed allele frequency changes are generally well predicted by accounting for the different genetic contributions of founders. Our results show that the genetic contribution of recent immigrants is substantial, with some large allele frequency shifts that otherwise may have been attributed to selection actually due to gene flow. We identified a few SNPs under directional short-term selection after appropriately accounting for gene flow. Using models that account for changes in population size, we partitioned the proportion of variance in allele frequency change through time. Observed allele frequency changes are primarily due to variation in survival and reproductive success, with gene flow making a smaller contribution. This study provides one of the most complete descriptions of short-term evolutionary change in allele frequencies in a natural population to date.

evolutionary biology

Detecting adaptive differentiation in structured populations with genomic data and common gardens

Adaptation in quantitative traits often occurs through subtle shifts in allele frequencies at many loci, a process called polygenic adaptation. While a number of methods have been developed to detect polygenic adaptation in human populations, we lack clear strategies for doing so in many other systems. In particular, there is an opportunity to develop new methods that leverage datasets with genomic data and common garden trait measurements to systematically detect the quantitative traits important for adaptation. Here, we develop methods that do just this, using principal components of the relatedness matrix to detect excess divergence consistent with polygenic adaptation and using a conditional test to control for confounding effects due to population structure. We apply these methods to inbred maize lines from the USDA germplasm pool and maize landraces from Europe. Ultimately, these methods can be applied to additional domesticated and wild species to give us a broader picture of the specific traits that contribute to adaptation and the overall importance of polygenic adaptation in shaping quantitative trait variation.

evolutionary biology

Reduced signal for polygenic adaptation of height in UK Biobank

Several recent papers have reported strong signals of selection on European polygenic height scores. These analyses used height effect estimates from the GIANT consortium and replication studies. Here, we describe a new analysis based on the the UK Biobank (UKB), a large, independent dataset. We find that the signals of selection using UKB effect-size estimates for height are strongly attenuated or absent. We also provide evidence that previous analyses were confounded by population stratification Therefore, the conclusion of strong polygenic adaptation now lacks support. Moreover, these discrepancies highlight (1) that methods for correcting for population stratification in GWAS may not always be sufficient for polygenic trait analyses, and (2) that claims of differences in polygenic scores between populations should be treated with caution until these issues are better understood.

evolutionary biology

Inferring Continuous and Discrete Population Genetic Structure Across Space

A classic problem in population genetics is the characterization of discrete population structure in the presence of continuous patterns of genetic differentiation. Especially when sampling is discontinuous, the use of clustering or assignment methods may incorrectly ascribe differentiation due to continuous processes (e.g., geographic isolation by distance) to discrete processes, such as geographic, ecological, or reproductive barriers between populations. This reflects a shortcoming of current methods for inferring and visualizing population structure when applied to genetic data deriving from geographically distributed populations. Here, we present a statistical framework for the simultaneous inference of continuous and discrete patterns of population structure. The method estimates ancestry proportions for each sample from a set of two-dimensional population layers, and, within each layer, estimates a rate at which relatedness decays with distance. This thereby explicitly addresses the \"clines versus clusters\" problem in modeling population genetic variation. The method produces useful descriptions of structure in genetic relatedness in situations where separated, geographically distributed populations interact, as after a range expansion or secondary contact. We demonstrate the utility of this approach using simulations and by applying it to empirical datasets of poplars and black bears in North America.\n\nAuthor summaryOne of the first steps in the analysis of genetic data, and a principal mission of biology, is to describe and categorize natural variation. A continuous pattern of differentiation (isolation by distance), where individuals found closer together in space are, on average, more genetically similar than individuals sampled farther apart, can confound attempts to categorize natural variation into groups. This is because current statistical methods for assigning individuals to discrete clusters cannot accommodate spatial patterns, and so are forced to use clusters to describe what is in fact continuous variation. As isolation by distance is common in nature, this is a substantial shortcoming of existing methods. In this study, we introduce a new statistical method for categorizing natural genetic variation - one that describes variation as a combination of continuous and discrete patterns. We demonstrate that this method works well and can capture patterns in population genomic data without resorting to splitting populations where they can be described by continuous patterns of variation.

evolutionary biology

Polygenic Adaptation has Impacted Multiple Anthropometric Traits

Our understanding of the genetic basis of human adaptation is biased toward loci of large pheno-typic effect. Genome wide association studies (GWAS) now enable the study of genetic adaptation in polygenic phenotypes. We test for polygenic adaptation among 187 world-wide human populations using polygenic scores constructed from GWAS of 34 complex traits. We identify signals of polygenic adaptation for anthropometric traits including height, infant head circumference (IHC), hip circumference and waist-to-hip ratio (WHR). Analysis of ancient DNA samples indicates that a north-south cline of height within Europe and and a west-east cline across Eurasia can be traced to selection for increased height in two late Pleistocene hunter gatherer populations living in western and west-central Eurasia. Our observation that IHC and WHR follow a latitudinal cline in Western Eurasia support the role of natural selection driving Bergmanns Rule in humans, consistent with thermoregulatory adaptation in response to latitudinal temperature variation.\n\nAuthors Note on Failure to ReplicateAfter this preprint was posted, the UK Biobank dataset was released, providing a new and open GWAS resource. When attempting to replicate the height selection results from this preprint using GWAS data from the UK Biobank, we discovered that we could not. In subsequent analyses, we determined that both the GIANT consortium height GWAS data, as well as another dataset that was used for replication, were impacted by stratification issues that created or at a minimum substantially inflated the height selection signals reported here. The results of this second investigation, written together with additional coauthors, have now been published (https://elifesciences.org/articles/39725 along with another paper by a separate group of authors, showing similar issues https://elifesciences.org/articles/39702). A preliminary investigation shows that the other non-height based results may suffer from similar issues. We stand by the theory and statistical methods reported in this paper, and the paper can be cited for these results. However, we have shown that the data on which the major empirical results were based are not sound, and so should be treated with caution until replicated.

evolutionary biology

Distinguishing among modes of convergent adaptation using population genomic data

Geographically separated populations can convergently adapt to the same selection pressure. Convergent evolution at the level of a gene may arise via three distinct modes. The selected alleles can (1) have multiple independent mutational origins, (2) be shared due to shared ancestral standing variation, or (3) spread throughout subpopulations via gene flow. We present a model-based, statistical approach that utilizes genomic data to detect cases of convergent adaptation at the genetic level, identify the loci involved and distinguish among these modes. To understand the impact of convergent positive selection on neutral diversity at linked loci, we make use of the fact that hitchhiking can be modeled as an increase in the variance in neutral allele frequencies around a selected site within a population. We build on coalescent theory to show how shared hitchhiking events between subpopulations act to increase covariance in allele frequencies between subpopulations at loci near the selected site, and extend this theory under different models of migration and selection on the same standing variation. We incorporate this hitchhiking effect into a multivariate normal model of allele frequencies that also accounts for population structure. Based on this theory, we present a composite-likelihood-based approach that utilizes genomic data to identify loci involved in convergence, and distinguishes among alternate modes of convergent adaptation. We illustrate our method on genome-wide polymorphism data from two distinct cases of convergent adaptation. First, we investigate the adaptation for copper toxicity tolerance in two populations of the common yellow monkey flower, Mimulus guttatus. We show that selection has occurred on an allele that has been standing in these populations prior to the onset of copper mining in this region. Lastly, we apply our method to data from four populations of the killifish, Fundulus heteroclitus, that show very rapid convergent adaptation for tolerance to industrial pollutants. Here, we identify a single locus at which both independent mutation events and selection on an allele shared via gene flow, either slightly before or during selection, play a role in adaptation across the species range.

evolutionary biology

Deconstructing isolation-by-distance: the genomic consequences of limited dispersal

Geographically limited dispersal can shape genetic population structure and result in a correlation between genetic and geographic distance, commonly called isolation-bydistance. Despite the prevalence of isolation-by-distance in nature, to date few studies have empirically demonstrated the processes that generate this pattern, largely because few populations have direct measures of individual dispersal and pedigree information. Intensive, long-term demographic studies and exhaustive genomic surveys in the Florida Scrub-Jay (Aphelocoma coerulescens) provide an excellent opportunity to investigate the influence of dispersal on genetic structure. Here, we used a panel of genome-wide SNPs and extensive pedigree information to explore the role of limited dispersal in shaping patterns of isolation-by-distance in both sexes, and at an exceedingly fine spatial scale (within ~10 km). Isolation-by-distance patterns were stronger in male-male and male-female comparisons than in female-female comparisons, consistent with observed differences in dispersal propensity between the sexes. Using the pedigree, we demonstrated how various genealogical relationships contribute to fine-scale isolation-by-distance. Simulations using field-observed distributions of male and female natal dispersal distances showed good agreement with the distribution of geographic distances between breeding individuals of different pedigree relationship classes. Furthermore, we extended Malecots theory of isolation-by-distance by building coalescent simulations parameterized by the observed dispersal curve, population density, and immigration rate, and showed how incorporating these extensions allows us to accurately reconstruct observed sex-specific isolation-by-distance patterns in autosomal and Z-linked SNPs. Therefore, patterns of fine-scale isolation-by-distance in the Florida Scrub-Jay can be well understood as a result of limited dispersal over contemporary timescales.\n\nAuthor SummaryDispersal is a fundamental component of the life history of most organisms and therefore influences many biological processes. Dispersal is particularly important in creating genetic structure on the landscape. We often observe a pattern of decreased genetic relatedness between individuals as geographic distances increases, or isolation-by-distance. This pattern is particularly pronounced in organisms with extremely short dispersal distances. Despite the ubiquity of isolation-by-distance patterns in nature, there are few examples that explicitly demonstrate how limited dispersal influences spatial genetic structure. Here we investigate the processes that result in spatial genetic structure using the Florida Scrub-Jay, a bird with extremely limited dispersal behavior and extensive genome-wide data. We take advantage of the long-term monitoring of a contiguous population of Florida Scrub-Jays, which has resulted in a detailed pedigree and measurements of dispersal for hundreds of individuals. We show how limited dispersal results in close genealogical relatives living closer together geographically, which generates a strong pattern of isolation-by-distance at an extremely small spatial scale (<10 km) in just a few generations. Given the detailed dispersal, pedigree, and genomic data, we can achieve a fairly complete understanding of how dispersal shapes patterns of genetic diversity over short spatial scales.

evolutionary biology