Search bioRxiv⌕ Search

Biology subjects

Irving-Pease, E. K.

Publications and source records attributed to Irving-Pease, E. K..

3 recordsLinked to original sources

The Selection Landscape and Genetic Legacy of Ancient Eurasians

The Holocene (beginning [~]12,000 years ago) encompassed some of the most significant changes in human evolution, with far-reaching consequences for the dietary, physical, and mental health of present-day populations. Using a dataset of >1600 imputed ancient genomes 1, we modelled the selection landscape during the transition from hunting and gathering, to farming and pastoralism across West Eurasia. We identify major selection signals related to metabolism, including that selection at the FADS cluster began earlier than previously reported, and that selection near the LCT locus predates the emergence of the lactase persistence allele by thousands of years. We also find strong selection in the HLA region, possibly due to increased exposure to pathogens during the Bronze Age. Using ancient individuals to infer local ancestry tracts in >400,000 samples from the UK Biobank, we identify widespread differences in the distribution of Mesolithic, Neolithic, and Bronze Age ancestries across Eurasia. By calculating ancestry-specific polygenic risk scores, we show that height differences between Northern and Southern Europe are associated with differential Steppe ancestry, rather than selection, and that risk alleles for mood-related phenotypes are enriched for Neolithic farmer ancestry, while risk alleles for diabetes and Alzheimers disease are enriched for Western Hunter-gatherer ancestry. Our results suggest that ancient selection and migration were major contributors to the distribution of phenotypic diversity in present-day Europeans.

evolutionary biology↗

Population Genomics of Stone Age Eurasia

Western Eurasia witnessed several large-scale human migrations during the Holocene1-5. To investigate the cross-continental impacts we shotgun-sequenced 317 primarily Mesolithic and Neolithic genomes from across Northern and Western Eurasia. These were imputed alongside published data to obtain diploid genotypes from >1,600 ancient humans. Our analyses revealed a Great Divide genomic boundary extending from the Black Sea to the Baltic. Mesolithic hunter-gatherers (HGs) were highly genetically differentiated east and west of this zone, and the impact of the neolithisation was equally disparate. Large-scale ancestry shifts occurred in the west as farming was introduced, including near-total replacements of HGs in many areas, whereas no substantial ancestry shifts happened east of the zone during the same period. Similarly, relatedness decreased in the west from the Neolithic transition onwards, while east of the Urals relatedness remained high until [~]4,000 BP, consistent with persistence of localised HG groups. The boundary dissolved when Yamnaya-related ancestry spread across western Eurasia around 5,000 BP resulting in a second major turnover that reached most parts of Europe within a 1,000-year span. The genetic origin and fate of the Yamnaya have remained elusive but we demonstrate that HGs from the Middle Don region contributed ancestry to them. Yamnaya-groups later admixed with individuals associated with the Globular Amphora Culture before expanding into Europe. Similar turnovers occurred in western Siberia, where we report new genomic data from a Neolithic steppe cline spanning the Siberian forest steppe to Lake Baikal. These prehistoric migrations had profound and lasting effects on the genetic diversity of Eurasian populations.

evolutionary biology↗

HAYSTAC: A Bayesian framework for robust and rapid species identification in high-throughput sequencing data

Identification of specific species in metagenomic samples is critical for several key applications, yet many tools available require large computational power and are often prone to false positive identifications. Here we describe High-AccuracY and Scalable Taxonomic Assignment of MetagenomiC data (HAYSTAC), which can estimate the probability that a specific taxon is present in a metagenome. HAYSTAC provides a user-friendly tool to construct databases, based on publicly available genomes, that are used for competitive reads mapping. It then uses a novel Bayesian framework to infer the abundance and statistical support for each species identification and provide per-read species classification. Unlike other methods, HAYSTAC is specifically designed to efficiently handle both ancient and modern DNA data, as well as incomplete reference databases, making it possible to run highly accurate hypothesis-driven analyses (i.e., assessing the presence of a specific species) on variably sized reference databases while dramatically improving processing speeds. We tested the performance and accuracy of HAYSTAC using simulated Illumina libraries, both with and without ancient DNA damage, and compared the results to other currently available methods (i.e., Kraken2/Bracken, KrakenUniq, MALT/HOPS, and Sigma). HAYSTAC identified fewer false positives than both Kraken2/Bracken, KrakenUniq and MALT in all simulations, and fewer than Sigma in simulations of ancient data. It uses less memory than Kraken2/Bracken, KrakenUniq as well as MALT both during database construction and sample analysis. Lastly, we used HAYSTAC to search for specific pathogens in two published ancient metagenomic datasets, demonstrating how it can be applied to empirical datasets. HAYSTAC is available from https://github.com/antonisdim/HAYSTAC Author summaryThe emerging field of paleo-metagenomics (i.e., metagenomics from ancient DNA) holds great promise for novel discoveries in fields as diverse as pathogen evolution and paleoenvironmental reconstruction. However, there is presently a lack of computational methods for species identification from microbial communities in both degraded and nondegraded DNA material. Here, we present "HAYSTAC", a user-friendly software package that implements a novel probabilistic model for species identification in metagenomic data obtained from both degraded and non-degraded DNA material. Through extensive benchmarking, we show that HAYSTAC can be used for accurately profiling the community composition, as well as for direct hypothesis testing for the presence of extremely low-abundance taxa, in complex metagenomic samples. After analysing simulated and publicly available datasets, HAYSTAC consistently produced the lowest number of false positive identifications during taxonomic profiling, produced robust results when databases of restricted size were used, and showed increased sensitivity for pathogen detection compared to other specialist methods. The newly proposed probabilistic model and software employed by HAYSTAC can have a substantial impact on the robust and rapid pathogen discovery in degraded/shallow sequenced metagenomic samples while optimising the use of computational resources.

bioinformatics↗