Search bioRxivSearch

Biology subjects

Thornton, K.

Publications and source records attributed to Thornton, K..

2 recordsLinked to original sources

Genetic architecture and selective sweeps after polygenic adaptation to distant trait optima

Understanding the genetic basis of phenotypic adaptation to changing environments is an essential goal of population and quantitative genetics. While technological advances now allow interrogation of genome-wide genotyping data in large panels, our understanding of the process of polygenic adaptation is still limited. To address this limitation, we use extensive forward-time simulation to explore the impacts of variation in demography, trait genetics, and selection on the rate and mode of adaptation and the resulting genetic architecture. We simulate a population adapting to an optimum shift, modeling sequence variation for 20 QTL for each of 12 different demographies for 100 different traits varying in the effect size distribution of new mutations, the strength of stabilizing selection, and the contribution of the genomic background. We then use random forest regression approaches to learn the relative importance of input parameters in determining a number of aspects of the process of adaptation including the speed of adaptation, the relative frequency of hard sweeps and sweeps from standing variation, or the final genetic architecture of the trait. We find that selective sweeps occur even for traits under relatively weak selection and where the genetic background explains most of the variation. Though most sweeps occur from variation segregating in the ancestral population, new mutations can be important for traits under strong stabilizing selection that undergo a large optimum shift. We also show that population bottlenecks and expansion impact overall genetic variation as well as the relative importance of sweeps from standing variation and the speed with which adaptation can occur. We then compare our results to two traits under selection during maize domestication, showing that our simulations qualitatively recapitulate differences between them. Overall, our results underscore the complex population genetics of individual loci in even relatively simple quantitative trait models, but provide a glimpse into the factors that drive this complexity and the potential of these approaches for understanding polygenic adaptation.\n\nAuthor summaryMany traits are controlled by a large number of genes, and environmental changes can lead to shifts in trait optima. How populations adapt to these shifts depends on a number of parameters including the genetic basis of the trait as well as population demography. We simulate a number of trait architectures and population histories to study the genetics of adaptation to distant trait optima. We find that selective sweeps occur even in traits under relatively weak selection and our machine learning analyses find that demography and the effect sizes of mutations have the largest influence on genetic variation after adaptation. Maize domestication is a well suited model for trait adaptation accompanied by demographic changes. We show how two example traits under a maize specific demography adapt to a distant optimum and demonstrate that polygenic adaptation is a well suited model for crop domestication even for traits with major effect loci.

evolutionary biology

Efficient pedigree recording for fast population genetics simulation

In this paper we describe how to efficiently record the entire genetic history of a population in forwards-time, individual-based population genetics simulations with arbitrary breeding models, population structure and demography. This approach dramatically reduces the computational burden of tracking individual genomes by allowing us to simulate only those loci that may affect reproduction (those having non-neutral variants). The genetic history of the population is recorded as a succinct tree sequence as introduced in the software package msprime, on which neutral mutations can be quickly placed afterwards. Recording the results of each breeding event requires storage that grows linearly with time, but there is a great deal of redundancy in this information. We solve this storage problem by providing an algorithm to quickly simplify a tree sequence by removing this irrelevant history for a given set of genomes. By periodically simplifying the history with respect to the extant population, we show that the total storage space required is modest and overall large efficiency gains can be made over classical forward-time simulations. We implement a general-purpose framework for recording and simplifying genealogical data, which can be used to make simulations of any population model more efficient. We modify two popular forwards-time simulation frameworks to use this new approach and observe efficiency gains in large, whole-genome simulations of one to two orders of magnitude. In addition to speed, our method for recording pedigrees has several advantages: (1) All marginal genealogies of the simulated individuals are recorded, rather than just genotypes. (2) A population of N individuals with M polymorphic sites can be stored in O(N log N + M) space, making it feasible to store a simulations entire final generation as well as its history. (3) A simulation can easily be initialized with a more efficient coalescent simulation of deep history. The software for recording and processing tree sequences is named tskit.\n\nAuthor SummarySexually reproducing organisms are related to the others in their species by the complex web of parent-offspring relationships that constitute the pedigree. In this paper, we describe a way to record all of these relationships, as well as how genetic material is passed down through the pedigree, during a forwards-time population genetic simulation. To make effective use of this information, we describe both efficient storage methods for this embellished pedigree as well as a way to remove all information that is irrelevant to the genetic history of a given set of individuals, which dramatically reduces the required amount of storage space. Storing this information allows us to produce whole-genome sequence from simulations of large populations in which we have not explicitly recorded new genomic mutations; we find that this results in computational run times of up to 50 times faster than simulations forced to explicitly carry along that information.

bioinformatics