Search bioRxivSearch

Biology subjects

Jay, F.

Publications and source records attributed to Jay, F..

2 recordsLinked to original sources

Joint ancestry inference reveals the landscape of archaic introgression in admixed populations

Studying the evolutionary history of archaic segments in recently admixed individuals requires inferring both continental and archaic ancestry in admixed genomes. Here, we present TRACTINATOR, the first deep-learning method for simultaneous inference of continental and archaic ancestry in admixed human genomes. The model combines SNP sequences, population allele-frequency information, and S* statistics to improve both inference tasks. By learning relationships between haplotypes and population allele frequencies, TRACTINATOR can generalize across genomic regions and even across different genomic datasets. We train our model using both real and synthetic data, and show that augmenting with synthetic data improves accuracy for both continental and archaic ancestry inference. Finally, we apply TRACTINATOR to admixed Latin American populations from the 1,000 Genomes Project, revealing how archaic ancestry is distributed within chromosomal segments of African, European and Indigenous American ancestry in Latin American individuals. For candidates of adaptive introgression, we also infer whether the archaic haplotype was introduced via European or Indigenous American ancestors.

bioinformatics

An ABC method for whole-genome sequence data: inferring paleolithic and neolithic human expansions

Species generally undergo a complex demographic history, consisting, in particular, of multiple changes in population size. Genome-wide sequencing data are potentially highly informative for reconstructing this demographic history. A crucial point is to extract the relevant information from these very large datasets. Here we designed an approach for inferring past demographic events from a moderate number of fully sequenced genomes. Our new approach uses Approximate Bayesian Computation (ABC), a simulation-based statistical framework that allows (i) identifying the best demographic scenario among several competing scenarios, and (ii) estimating the best-fitting parameters under the chosen scenario. ABC relies on the computation of summary statistics. Using a cross-validation approach, we showed that statistics such as the lengths of haplotypes shared between individuals, or the decay of linkage disequilibrium with distance, can be combined with classical statistics (eg heterozygosity, Tajimas D) to accurately infer complex demographic scenarios including bottlenecks and expansion periods. We also demonstrated the importance of simultaneously estimating the genotyping error rate. Applying our method on genome-wide human-sequence databases, we finally showed that a model consisting in a bottleneck followed by a Paleolithic and a Neolithic expansion was the most relevant for Eurasian populations.

evolutionary biology