Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.04.17.589597

The Genetic Origin of the Indo-Europeans

Abstract

The Yamnaya archaeological complex appeared around 3300BCE across the steppes north of the Black and Caspian Seas, and by 3000BCE reached its maximal extent from Hungary in the west to Kazakhstan in the east. To localize the ancestral and geographical origins of the Yamnaya among the diverse Eneolithic people that preceded them, we studied ancient DNA data from 428 individuals of which 299 are reported for the first time, demonstrating three previously unknown Eneolithic genetic clines. First, a "Caucasus-Lower Volga" (CLV) Cline suffused with Caucasus hunter-gatherer (CHG) ancestry extended between a Caucasus Neolithic southern end in Neolithic Armenia, and a steppe northern end in Berezhnovka in the Lower Volga. Bidirectional gene flow across the CLV cline created admixed intermediate populations in both the north Caucasus, such as the Maikop people, and on the steppe, such as those at the site of Remontnoye north of the Manych depression. CLV people also helped form two major riverine clines by admixing with distinct groups of European hunter-gatherers. A "Volga Cline" was formed as Lower Volga people mixed with upriver populations that had more Eastern hunter-gatherer (EHG) ancestry, creating genetically hyper-variable populations as at Khvalynsk in the Middle Volga. A "Dnipro Cline" was formed as CLV people bearing both Caucasus Neolithic and Lower Volga ancestry moved west and acquired Ukraine Neolithic hunter-gatherer (UNHG) ancestry to establish the population of the Serednii Stih culture from which the direct ancestors of the Yamnaya themselves were formed around 4000BCE. This population grew rapidly after 3750-3350BCE, precipitating the expansion of people of the Yamnaya culture who totally displaced previous groups on the Volga and further east, while admixing with more sedentary groups in the west. CLV cline people with Lower Volga ancestry contributed four fifths of the ancestry of the Yamnaya, but also, entering Anatolia from the east, contributed at least a tenth of the ancestry of Bronze Age Central Anatolians, where the Hittite language, related to the Indo-European languages spread by the Yamnaya, was spoken. We thus propose that the final unity of the speakers of the "Proto-Indo-Anatolian" ancestral language of both Anatolian and Indo-European languages can be traced to CLV cline people sometime between 4400-4000 BCE. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=115 SRC="FIGDIR/small/589597v1_ufig1.gif" ALT="Figure 1"> View larger version (91K): org.highwire.dtl.DTLVardef@127de98org.highwire.dtl.DTLVardef@87010aorg.highwire.dtl.DTLVardef@1554627org.highwire.dtl.DTLVardef@170dd63_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOSummary Figure:C_FLOATNO The origin of Indo-Anatolian and Indo-European languages. Genetic reconstruction of the ancestry of Pontic-Caspian steppe and West Asian populations points to the North Caucasus-Lower Volga area as the homeland of Indo-Anatolian languages and to the Serednii Stih archaeological culture of the Dnipro-Don area as the homeland of Indo-European languages. The Caucasus-Lower Volga people had diverse distal roots, estimated using the qpAdm software on the left barplot, as Caucasus hunter-gatherer (purple), Central Asian (red), Eastern hunter-gatherer (pink), and West Asian Neolithic (green). Caucasus-Lower Volga expansions, estimated using qpAdm on the right barplot as disseminated Caucasus Neolithic (blue)-Lower Volga Eneolithic (orange) proximal ancestries, mixing with the inhabitants of the North Pontic region (yellow), Volga region (yellow), and West Asia (green). C_FIG

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lazaridis, I., Patterson, N., Anthony, D., Vyazov, L., Fournier, R., Ringbauer, H., Olalde, I., Khokhlov, A. A., Kitov, E. P., Shishlina, N. I., Ailincai, S. C., Agapov, D. S., Agapov, S. A., Batieva, E., Bauyrzhan, B., Bereczki, Z., Buzhilova, A., Changmai, P., Chizhevsky, A. A., Ciobanu, I., Constantinescu, M., Csanyi, M., Dani, J., Dashkovskiy, P. K., Evinger, S., Faifert, A., Flegontov, P. N., Frinculeasa, A., Frinculeasa, M. N., Hajdu, T., Higham, T., Jarosz, P., Jelinek, P., Khartanovich, V. I., Kirginekov, E. N., Kiss, V., Kitova, A., Kiyashko, A. V., Koledin, J., Korolev, A., Kosintsev. 2024-04-18. The Genetic Origin of the Indo-Europeans. https://doi.org/10.1101/2024.04.17.589597

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Generation of a transgenic cephalopod

Coleoid cephalopods (cuttlefish, octopus, and squid) are marine mollusks with elaborate nervous systems that support a diverse repertoire of complex behaviors. These include the neural control of the color, pattern, and texture of the skin, facilitating both adaptive camouflage and innate patterning that may reflect internal state. The development of transgenic cephalopods expressing fluorescent proteins, optogenetic actuators, and reporters of neural activity would contribute a new and important technology to cephalopod biology. The generation of transgenic cephalopods, however, has remained a major challenge. Here, we report the development of stable transgenic dwarf cuttlefish (Ascarosepion bandense) expressing ubiquitous nuclear-localized mScarlet, a red fluorescent protein. We evaluated multiple strategies for transgenesis, and established cuttlefish lines using both CRISPR and the transposons Sleeping Beauty and Minos. The stable expression of transgenes enabled live imaging of cell dynamics during embryonic development. The Minos transposon emerged as the most efficient transgenesis strategy and is adaptable to promoters and transgenes of choice. These strategies now enable the generation of diverse genetic tools for mechanistic studies of cephalopod biology.

genetics↗

Large language model-based bibliometric evaluation of population descriptors in human genetics

As the use of population descriptors such as race, ethnicity, and ancestry have become increasingly common in modern genetics research, there have been growing calls to critically examine their use. Most notably, in 2023, the National Academies of Science, Engineering, and Medicine (NASEM) published a report titled Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field, which included eight specific and actionable recommendations for researchers to implement the ethical and accurate use of population descriptors in genetic research. Here, we use the 2023 NASEM report as a benchmark to analyze the use of population descriptors in genome-wide association studies (GWAS). We develop a general toolkit for large language model-based bibliometrics, operationalize the report's recommendations into an evaluation framework, and apply this framework to evaluate all 4,007 papers from the GWAS Catalog published between 2007 and 2025 with full text available on PubMedCentral. We find significant improvements in adherence to NASEM report recommendations over time. However, most improvements predate the publication of the NASEM report itself, suggesting the report functioned primarily as a synthesis of existing best practices rather than a catalyst for change. We conclude by highlighting opportunities for growth in the field of human genetics.

genetics↗

Mitigating biases of rescaling in forward-in-time population genetic simulations

Forward-in-time population genetic simulations are widely used in evolutionary analyses, but simulating large populations and long genomic regions remains computationally demanding. To reduce this cost, parameter rescaling is widely employed, in which the original evolutionary process is approximated by one with a smaller population size and fewer generations. Recently, several studies using the SLiM simulator have raised concerns about the accuracy of this rescaling approach. In this study, we show that many of the biases reported in these studies can be mitigated by using a different simulation algorithm. These results reveal that the accuracy of parameter rescaling depends on how well the simulation algorithm preserves diffusion-limit properties under rescaling.

genetics↗