Search bioRxiv⌕ Search

Biology subjects

Smith, O. S.

Publications and source records attributed to Smith, O. S..

3 recordsLinked to original sources

Representation in genetic studies affects inference about genetic architecture

Knowledge of a trait's "genetic architecture," namely the joint distribution of allele frequencies of causal variants and the direction and magnitude of their effects, is essential to understanding its evolution and underlying biology. Inferences about genetic architecture are based on data collected in different ways in cohorts recruited through heterogeneous mechanisms. As a result, cohorts differ in genotype, environment, and trait distributions. For example, the UK Biobank (UKB) was designed for broad population representation, whereas FinnGen drew extensively from clinical registries enriched for diagnosed health conditions. Here, we asked whether representation in genetic studies influences inferences about genetic architectures. Using GWAS data from the UKB, All of Us (AoU), and FinnGen, we compared several summaries of genetic architecture. We find systematic differences in some summaries, such as SNP heritability, which were often (in 6/13 traits) significantly lower in AoU than in the UKB (and never the reverse), even when matching samples such that they have similar genetic ancestry compositions. This result aligns with other recent evidence that biobanks enriched for diagnosed health conditions, also sometimes characterized by less-standardized phenotyping, have lower heritability than population-based biobanks. We highlight a second case, where a summary of genetic architecture varies considerably but not systematically across traits and biobanks. Such is the case for the mean direction of allelic effects ("sign bias"). For example, 72% of rare minor alleles affecting type 2 diabetes risk are inferred to be risk-increasing based on AoU data, while nearly all (>99%) are inferred to be risk-increasing based on UKB data. We hypothesize that the inferred sign bias is heavily influenced by the skewness of the trait distribution in the study and otherwise largely independent of other study or trait characteristics, including whether the trait is binary or quantitative. We provide strong support for this hypothesis through simulations and data from the three biobanks: the variation in inferred sign bias for rare minor alleles across traits and biobanks is explained remarkably well (82% and 97% of variance explained for trait-associated SNPs and randomly selected SNPs, respectively) solely by the trait's skewness in the biobank, with residual biobank-specificity explaining little. Our findings suggest that inferences about the map between genetic and trait variation can depend on study design and participation in genetic studies in surprising ways.

genomics↗

A Litmus Test for Confounding in Polygenic Scores

Polygenic scores (PGSs) are being rapidly adopted for trait prediction in the clinic and beyond. PGSs are often thought of as capturing the direct genetic effect of one's genotype on one's phenotype. However, because PGSs are constructed from population-level associations, they are influenced by factors other than direct genetic effects, including stratification, assortative mating, and dynastic effects ("SAD effects"). Our interpretation and application of PGSs may hinge on the relative influence of SAD effects, since they may often be environmentally or culturally mediated. We developed a method to measure these influences, Partitioning Genetic Scores Using Siblings (PGSUS, pron. "Pegasus"). PGSUS leverages a comparison of a PGS of interest based on a standard GWAS with a PGS based on a sibling GWAS--which is largely immune to SAD effects--to partition variance in a PGS (in a given sample) into components due to direct effects, SAD effects, and their covariance. Using PGSUS, we found that in many cases direct genetic effects contribute relatively little to PGS variation--most pronouncedly so in PGSs for social or behavioral traits, such as educational attainment or neuroticism. PGSUS further breaks down variance components by axes of genetic ancestry, allowing for a nuanced interpretation of SAD effects. In particular, PGSUS can detect stratification along major axes of ancestry as well as SAD variance that is "isotropic" with respect to axes of ancestry. Applying PGSUS, we found evidence of stratification in PGSs constructed using large meta-analyses of height as well as in multiple PGSs constructed using the UK Biobank. We show that a given PGS can suffer from stratification along a major axis of ancestry in one sample but not in another (for example, in comparisons of prediction in samples from contemporary vs. ancient DNA samples). We further show that when axes of stratification are shared between GWAS and prediction samples, stratification can both aid and impede phenotypic prediction accuracy. In summary, PGSUS offers advances in interpretation towards more informed application of polygenic scores.

genomics↗

1,000 ancient genomes uncover 10,000 years of natural selection in Europe

Ancient DNA has revolutionized our understanding of human population history. However, its potential to examine how rapid cultural evolution to new lifestyles may have driven biological adaptation has not been met, largely due to limited sample sizes. We assembled genome-wide data from 1,291 individuals from Europe over 10,000 years, providing a dataset that is large enough to resolve the timing of selection into the Neolithic, Bronze Age, and Historical periods. We identified 25 genetic loci with rapid changes in frequency during these periods, a majority of which were previously undetected. Signals specific to the Neolithic transition are associated with body weight, diet, and lipid metabolism-related phenotypes. They also include immune phenotypes, most notably a locus that confers immunity to Salmonella infection at a time when ancient Salmonella genomes have been shown to adapt to human hosts, thus providing a possible example of human-pathogen co-evolution. In the Bronze Age, selection signals are enriched near genes involved in pigmentation and immune-related traits, including at a key human protein interactor of SARS-CoV-2. Only in the Historical period do the selection candidates we detect largely mirror previously-reported signals, highlighting how the statistical power of previous studies was limited to the last few millennia. The Historical period also has multiple signals associated with vitamin D binding, providing evidence that lactase persistence may have been part of an oligogenic adaptation for efficient calcium uptake and challenging the theory that its adaptive value lies only in facilitating caloric supplementation during times of scarcity. Finally, we detect selection on complex traits in all three periods, including selection favoring variants that reduce body weight in the Neolithic. In the Historical period, we detect selection favoring variants that increase risk for cardiovascular disease plausibly reflecting selection for a more active inflammatory response that would have been adaptive in the face of increased infectious disease exposure. Our results provide an evolutionary rationale for the high prevalence of these deadly diseases in modern societies today and highlight the unique power of ancient DNA in elucidating biological change that accompanied the profound cultural transformations of recent human history.

genomics↗