Search bioRxiv⌕ Search

Biology subjects

Rosengren, A.

Publications and source records attributed to Rosengren, A..

2 recordsLinked to original sources

The Selection Landscape and Genetic Legacy of Ancient Eurasians

The Holocene (beginning [~]12,000 years ago) encompassed some of the most significant changes in human evolution, with far-reaching consequences for the dietary, physical, and mental health of present-day populations. Using a dataset of >1600 imputed ancient genomes 1, we modelled the selection landscape during the transition from hunting and gathering, to farming and pastoralism across West Eurasia. We identify major selection signals related to metabolism, including that selection at the FADS cluster began earlier than previously reported, and that selection near the LCT locus predates the emergence of the lactase persistence allele by thousands of years. We also find strong selection in the HLA region, possibly due to increased exposure to pathogens during the Bronze Age. Using ancient individuals to infer local ancestry tracts in >400,000 samples from the UK Biobank, we identify widespread differences in the distribution of Mesolithic, Neolithic, and Bronze Age ancestries across Eurasia. By calculating ancestry-specific polygenic risk scores, we show that height differences between Northern and Southern Europe are associated with differential Steppe ancestry, rather than selection, and that risk alleles for mood-related phenotypes are enriched for Neolithic farmer ancestry, while risk alleles for diabetes and Alzheimers disease are enriched for Western Hunter-gatherer ancestry. Our results suggest that ancient selection and migration were major contributors to the distribution of phenotypic diversity in present-day Europeans.

evolutionary biology↗

Accuracy of haplotype estimation and whole genome imputation affects complex trait analyses in complex biobanks

Sample recruitment for research consortia, hospitals, biobanks, and personal genomics companies span years, necessitating genotyping in batches, using different technologies. As marker content on genotyping arrays varies systematically, integrating such datasets is non-trivial and its impact on haplotype estimation (phasing) and whole genome imputation, necessary steps for complex trait analysis, remains under-evaluated. Using the iPSYCH consortium dataset, comprising 130,438 individuals, genotyped in two stages, on different arrays, we evaluated phasing and imputation performance across multiple phasing methods and data integration protocols. While phasing accuracy varied both by choice of method and data integration protocol, imputation accuracy varied mostly between data integration protocols. We demonstrate an attenuation in imputation accuracy within samples of non-European origin, highlighting challenges to studying complex traits in diverse populations. Finally, imputation errors can modestly bias association tests and reduce predictive utility of polygenic scores. This is the largest, most comprehensive comparison of data integration approaches in the context of a large psychiatric biobank.

bioinformatics↗