Search bioRxiv⌕ Search

Biology subjects

Cousins, T.

Publications and source records attributed to Cousins, T..

6 recordsLinked to original sources

Deep coalescent history of the hominin lineage

Coalescent-based methods are widely used to infer population size histories, but existing analyses have limited resolution at deep time scales (>2 million years ago). Here, using new high-quality telomere-to-telomere genome assemblies, we extend the scope of such inference and reanalyse an ancient peak in human effective population size around 3-6 million years ago, showing that coalescent-based inference can reach much further into the past than previously thought. Furthermore, peaks at approximately the same time as in humans are also observed in chimpanzee and bonobo, but not in gorilla or orangutans. We show that this pattern is unlikely to be an artefact of model violations and discuss its potential implications for understanding hominin evolutionary history. It is known that such peaks can arise from ancestral genetic structure, and consequently we suggest that a parsimonious explanation for the contemporaneous peaks in humans, chimpanzees and bonobos would be a period of complex speciation before final separation of Homo from Pan, in contrast to a scenario where these lineages diverge cleanly from a panmictic common ancestor.

genetics↗

Avoidable false PSMC population size peaks occur across numerous studies

Inferring historical population sizes is key to identify drivers of ecological and evolutionary change, and crucial to predict the future of species on our rapidly changing planet. The pairwise sequentially Markovian coalescent (PSMC) method provided a revolutionary framework to reconstruct species demographic histories over millions of years based on the genome sequence of a single individual 1. Here, we detected and solved a common artifact in PSMC and related methods: recent population peaks followed by population collapses. Combining real and simulated genomes, we show that these peaks do not represent true population dynamics. Instead, ill-set default parameters cause false peaks in our own and published data, which can be avoided by adjusted parameter settings. Furthermore, we show that certain population structure changes can cause similar patterns. Newer methods like Beta-PSMC perform better, but do not always avoid this artifact. Our results suggest testing multiple parameters before interpreting recent population peaks followed by collapses, and call for the development of robust methods.

evolutionary biology↗

A structured coalescent model reveals deep ancestral structure shared by all modern humans

1Understanding the series of admixture events and population size history leading to modern humans is central to human evolutionary genetics. Using a coalescence-based hidden Markov model, we present evidence for an extended period of structure in the history of all modern humans, in which two ancestral populations that diverged [~]1.5 million years ago came together in an admixture event [~]300 thousand years ago, in a ratio of [~]80:20 percent. Immediately after their divergence, we detect a strong bottleneck in the major ancestral population. We inferred regions of the present-day genome derived from each ancestral population, finding that material from the minority correlates strongly with distance to coding sequence, suggesting it was deleterious against the majority background. Moreover, we found a strong correlation between regions of majority ancestry and human-Neanderthal or human-Denisovan divergence, suggesting the majority population was also ancestral to those archaic humans.

genetics↗

Accurate inference of population history in the presence of background selection

1All published methods for learning about demographic history make the simplifying assumption that the genome evolves neutrally, and do not seek to account for the effects of natural selection on patterns of variation. This is a major concern, as ample work has demonstrated the pervasive effects of natural selection and in particular background selection (BGS) on patterns of genetic variation in diverse species. Simulations and theoretical work have shown that methods to infer changes in effective population size over time (Ne(t)) become increasingly inaccurate as the strength of linked selection increases. Here, we introduce an extension to the Pairwise Sequentially Markovian Coalescent (PSMC) algorithm, PSMC+, which explicitly co-models demographic history and natural selection. We benchmark our method using forward-in-time simulations with BGS and find that our approach improves the accuracy of effective population size inference. Leveraging a high resolution map of BGS in humans, we infer considerable changes in the magnitude of inferred effective population size relative to previous reports. Finally, we separately infer Ne(t) on the X chromosome and on the autosomes in diverse great apes without making a correction for selection, and find that the inferred ratio fluctuates substantially through time in a way that differs across species, showing that uncorrected selection may be an important driver of signals of genetic difference on the X chromosome and autosomes.

genetics↗

Lepidoptera genomics based on 88 chromosomal reference sequences informs population genetic parameters for conservation

Butterflies and moths (Lepidoptera) are one of the most ecologically diverse and speciose insect orders, with more than 157,000 described species. However, the abundance and diversity of Lepidoptera are declining worldwide at an alarming rate. As few Lepidoptera are explicitly recognised as at risk globally, the need for conservation is neither mandated nor well-evidenced. Large-scale biodiversity genomics projects that take advantage of the latest developments in long-read sequencing technologies offer a valuable source of information. We here present a comprehensive, reference-free, whole-genome, multiple sequence alignment of 88 species of Lepidoptera. We show that the accuracy and quality of the alignment is influenced by the contiguity of the reference genomes analysed. We explored genomic signatures that might indicate conservation concern in these species. In our dataset, which is largely from Britain, many species, in particular moths, display low heterozygosity and a high level of inbreeding, reflected in medium (0.1 - 1 Mb) and long (> 1 Mb) runs of homozygosity. Many species with low inbreeding display a higher masked load, estimated from the sum of rejected substitution scores at heterozygous sites. Our study shows that the analysis of a single diploid genome in a comparative phylogenetic context can provide relevant genetic information to prioritise species for future conservation investigation, particularly for those with an unknown conservation status.

genomics↗