Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.04.03.646226

A high-resolution genomic study of the Pama-Nyungan speaking Yolngu people of northeast Arnhem Land, Australia.

Abstract

ObjectivesAbout 300 Aboriginal languages were spoken in Australia. These were classified into two groups: Pama-Nyungan (PN), comprised of one language Family, and Non-Pama-Nyungan (NPN) with more than 20 language Families. The Yolngu people belong to the larger PN Family and live in Arnhem Land in northern Australia. They are surrounded by groups who speak NPN languages. This study, using nuclear genomic and mitochondrial DNA data, was undertaken to shed light on the origins of the Yolngu people and their language. The nuclear genomic sequences of Yolngu people were compared to those of other Indigenous Australians, as well as Papuan, African, East Asian and European people. Materials and methodsWith the agreement of Indigenous participants, samples were collected from 13 Yolngu individuals and 4 people from neighbouring NPN speakers and their nuclear genomes sequenced to a 30lil coverage. Using the short-read DNA BGISEQ-500 technology, these sequences were mapped to a reference genome and identified [~]24.86 million Single Nucleotide Variants (SNVs). The Yolngu SNVs were then compared to those of 36 individuals from 10 other Indigenous populations/locations across Australia and four worldwide populations using multidimensional scaling, population structure, F3 statistics and phylogenetic analyses. ResultsUsing the above methods, we infer that Yolngu speakers are closely related to neighbouring NPN speakers, followed by the Weipa population. No European or East Asian admixture was detected in the genomes of the Yolngu speakers studied here, which contrasts with the genomes of many other PN speakers that have been studied. Our results show that Yolngu speakers are more closely related to other PN speakers in the northeast of Australia than to those in central and western Australia studied here. Yolngu and the other Australian populations from this study share Papuans as an out-group. DiscussionThe study presented here provides an account of the nuclear and mitochondrial genomic diversity within the PN Yolngu Aboriginal population. The results show the Yolngu sample and their NPN neighbours have a strong genetic relationship. They also offer evidence of ancestral links between the Yolngu and PN-speaking populations in Cape York. From earlier fingerprint studies, consistent with the genomic results shown here, we suggest that there was a movement of people from the east into northeast Arnhem Land, associated with the flooding of the Sahul Shelf, and that this occurred between about 11 Kya and 8 Kya ago. Several Yolngu myths point to such a movement. It is suggested that the spread of the PN language or its speakers may have influenced the population structure of the Yolngu. Further genomic studies, with larger samples, of populations to the east of the Yolngu around the Gulf of Carpentaria into Cape York are required to test this hypothesis. Our results imply that PN did not spread with the movement of people across the continent, rather, the PN languages diffused among the different populations. It seems clear that the languages dispersed and not the people. The low level of relatedness detected between the Yolngu people and the people of the central arid desert of Australia suggests a long period of separation with different patterns of migration. Beyond Australia, Yolngu are most closely related to the Papuan people of New Guinea.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

White, N., Kumar, M., Lambert, D.. 2025-04-03. A high-resolution genomic study of the Pama-Nyungan speaking Yolngu people of northeast Arnhem Land, Australia.. https://doi.org/10.1101/2025.04.03.646226

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Utilizing single-cell data for per-cell type eQTL mapping in the human pancreas

Aims/hypothesis The human pancreas is a central organ for metabolic regulation that is comprised of diverse cell types that uniquely contribute to its function. Previous studies have performed expression quantitative trail loci (eQTL) discovery in either whole pancreas or in pancreatic islets, but due to differences between pancreatic cell types, this approach does not reveal cell type-specific effects. In this study, we sought to either implicate the cell type of action for known eQTLs or identify new eQTLs that may have been masked in bulk studies by performing eQTL discovery in individual pancreatic cell types. Methods We clustered 153,018 single-cell RNA sequencing (scRNA-seq) data from 71 pancreatic islet donors from the Human Pancreas Analysis Program (HPAP). We performed eQTL discovery in six pancreatic cell types using this resource directly. We further utilized this single cell resource as a reference to deconvolute bulk pancreatic RNA sequencing data from 305 Genotype Tissue Expression (GTEx) project donors and performed eQTL discovery in four pancreatic cell types. Finally, we performed fine-mapping and co-localization of pancreatic cell type eQTLs with metabolic GWAS to connect our findings to metabolic disease risk. Results From analyzing 71 individuals with single cell profiles, we identified 112 unique eGenes across six pancreatic cell types, 99 of which had been identified previously and 13 unique to this study. From the deconvoluted eQTLs, we identified 3,134 unique eGenes across four pancreatic cell types, 116 of which were unique to our study. Fine-mapping and co-localization of eQTLs with metabolic GWAS yielded key leads that warrant further investigation, such as the association of rs2168101 with LMO1 expression in alpha cells. Conclusions/interpretation We identified new signals that were previously not found in bulk pancreatic eQTL studies and potential cell type of action for several signals that were identified previously. Although there are limitations to the power, and therefore, discoverability of this study, it provides insights into how individual pancreatic cells differently contribute to metabolic disease.

genetics↗

Snurportin-1 maintains muscle niche integrity and myogenic progenitor homeostasis

Loss-of-function variants in SNUPN, encoding the nuclear import factor Snurportin-1 (SPN1) required for spliceosomal small nuclear ribonucleoprotein (snRNP) transport, cause a recently described form of limb-girdle muscular dystrophy (LGMD). However, the role of SPN1 in skeletal muscle homeostasis remains poorly understood, in part due to the lack of a suitable in vivo model. Here, we generated a zebrafish snupn loss-of-function model that recapitulates key features of the skeletal muscle phenotype observed in patients. Mutant larvae developed severe locomotor impairment by 6 days post-fertilization (dpf), accompanied by sarcomeric disorganization and impaired muscle fiber integrity. Transcriptomic profiling at 6 dpf revealed widespread alternative splicing and transcriptional dysregulation, with prominent alterations in extracellular matrix and basement membrane components, together with upregulation of stress- and inflammation-associated genes. Notably, these late-stage abnormalities were preceded by disruption of the muscle progenitor population at 2 dpf, with reduced Pax7 progenitor abundance and myogenic gene expression together with altered muscle differentiation and organization. Together, these findings identify SPN1 as a key regulator of skeletal muscle homeostasis linking RNA processing to extracellular niche integrity and myogenic progenitor maintenance. This zebrafish model provides an in vivo platform for dissecting LGMD-associated disease mechanisms and developing therapeutic strategies aimed at restoring muscle function and regenerative capacity.

genetics↗

MOD-scTWAS: Leveraging gene co-expression for single-cell transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) provide an effective framework for identifying genes associated with complex traits. Population-scale single-cell transcriptomic data enable genetically regulated expression (GReX) prediction and TWAS analyses at cell-type resolution, but the predictive performance of existing single-cell TWAS methods remains limited. Here, we develop MOD-scTWAS, a module-based method that jointly models GReX for genes within co-expression modules to borrow information across genes. Starting from a generative model for single-cell gene expression, MOD-scTWAS accounts for the heteroscedasticity and cross-gene correlation of individual-level pseudobulk expression in joint GReX prediction. In cross-validation analyses of the OneK1K dataset, MOD-scTWAS achieved higher mean GReX prediction accuracy than scTWAS across all 14 cell types and increased the number of imputable genes. When applied to TWAS analyses of UK Biobank quantitative hematological traits, MOD-scTWAS identified more significant cell type-gene-trait associations than scTWAS. These results demonstrate the potential of leveraging gene co-expression through joint modeling to improve cell-type-specific GReX prediction and TWAS discovery.

genetics↗