Search bioRxiv⌕ Search

bioRxiv · 10.1101/2025.01.23.634620

The central limit theorem for the number of mutations in the genealogy of a sample from a large population

Abstract

The number K of mutations identifiable in a sample of n sequences from a large population is one of the most important summary statistics in population genetics and is ubiquitous in the analysis of DNA sequence data. K can be expressed as the sum of n -1 independent geometric random variables. Consequently, its probability generating function was established long ago, yielding its well-known expectation and variance. However, the statistical properties of K is much less understood than those of the number of distinct alleles in a sample. This paper demonstrates that the central limit theorem holds for K, implying that K follows approximately a normal distribution when a large sample is drawn from a population evolving according to the Wright-Fisher model with a constant effective size, or according to the constant-in-state model, which allows population sizes to vary independently but bounded uniformly across different states of the coalescent process. Additionally, the skewness and kurtosis of K are derived, confirming that K has asymptotically the same skewness and kurtosis as a normal distribution. Furthermore, skewness converges at speed [Formula] and while kurtosis at speed 1 /ln n. Despite the overall convergence speed to normality is relatively slow, the distribution of K for a modest sample size is already not too far from normality, therefore the asymptotic normality may be sufficient for certain applications when the sample size is large enough.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fu, Y.-X.. 2025-01-26. The central limit theorem for the number of mutations in the genealogy of a sample from a large population. https://doi.org/10.1101/2025.01.23.634620

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Utilizing single-cell data for per-cell type eQTL mapping in the human pancreas

Aims/hypothesis The human pancreas is a central organ for metabolic regulation that is comprised of diverse cell types that uniquely contribute to its function. Previous studies have performed expression quantitative trail loci (eQTL) discovery in either whole pancreas or in pancreatic islets, but due to differences between pancreatic cell types, this approach does not reveal cell type-specific effects. In this study, we sought to either implicate the cell type of action for known eQTLs or identify new eQTLs that may have been masked in bulk studies by performing eQTL discovery in individual pancreatic cell types. Methods We clustered 153,018 single-cell RNA sequencing (scRNA-seq) data from 71 pancreatic islet donors from the Human Pancreas Analysis Program (HPAP). We performed eQTL discovery in six pancreatic cell types using this resource directly. We further utilized this single cell resource as a reference to deconvolute bulk pancreatic RNA sequencing data from 305 Genotype Tissue Expression (GTEx) project donors and performed eQTL discovery in four pancreatic cell types. Finally, we performed fine-mapping and co-localization of pancreatic cell type eQTLs with metabolic GWAS to connect our findings to metabolic disease risk. Results From analyzing 71 individuals with single cell profiles, we identified 112 unique eGenes across six pancreatic cell types, 99 of which had been identified previously and 13 unique to this study. From the deconvoluted eQTLs, we identified 3,134 unique eGenes across four pancreatic cell types, 116 of which were unique to our study. Fine-mapping and co-localization of eQTLs with metabolic GWAS yielded key leads that warrant further investigation, such as the association of rs2168101 with LMO1 expression in alpha cells. Conclusions/interpretation We identified new signals that were previously not found in bulk pancreatic eQTL studies and potential cell type of action for several signals that were identified previously. Although there are limitations to the power, and therefore, discoverability of this study, it provides insights into how individual pancreatic cells differently contribute to metabolic disease.

genetics↗

MOD-scTWAS: Leveraging gene co-expression for single-cell transcriptome-wide association studies

Transcriptome-wide association studies (TWAS) provide an effective framework for identifying genes associated with complex traits. Population-scale single-cell transcriptomic data enable genetically regulated expression (GReX) prediction and TWAS analyses at cell-type resolution, but the predictive performance of existing single-cell TWAS methods remains limited. Here, we develop MOD-scTWAS, a module-based method that jointly models GReX for genes within co-expression modules to borrow information across genes. Starting from a generative model for single-cell gene expression, MOD-scTWAS accounts for the heteroscedasticity and cross-gene correlation of individual-level pseudobulk expression in joint GReX prediction. In cross-validation analyses of the OneK1K dataset, MOD-scTWAS achieved higher mean GReX prediction accuracy than scTWAS across all 14 cell types and increased the number of imputable genes. When applied to TWAS analyses of UK Biobank quantitative hematological traits, MOD-scTWAS identified more significant cell type-gene-trait associations than scTWAS. These results demonstrate the potential of leveraging gene co-expression through joint modeling to improve cell-type-specific GReX prediction and TWAS discovery.

genetics↗

Generation of a transgenic cephalopod

Coleoid cephalopods (cuttlefish, octopus, and squid) are marine mollusks with elaborate nervous systems that support a diverse repertoire of complex behaviors. These include the neural control of the color, pattern, and texture of the skin, facilitating both adaptive camouflage and innate patterning that may reflect internal state. The development of transgenic cephalopods expressing fluorescent proteins, optogenetic actuators, and reporters of neural activity would contribute a new and important technology to cephalopod biology. The generation of transgenic cephalopods, however, has remained a major challenge. Here, we report the development of stable transgenic dwarf cuttlefish (Ascarosepion bandense) expressing ubiquitous nuclear-localized mScarlet, a red fluorescent protein. We evaluated multiple strategies for transgenesis, and established cuttlefish lines using both CRISPR and the transposons Sleeping Beauty and Minos. The stable expression of transgenes enabled live imaging of cell dynamics during embryonic development. The Minos transposon emerged as the most efficient transgenesis strategy and is adaptable to promoters and transgenes of choice. These strategies now enable the generation of diverse genetic tools for mechanistic studies of cephalopod biology.

genetics↗