Search bioRxivSearch

Biology subjects

Shpak, M.

Publications and source records attributed to Shpak, M..

2 recordsLinked to original sources

The Moran coalescent in a discrete one-dimensional spatial model

Among many organisms, offspring are constrained to occur at sites adjacent to their parents. This applies to plants and animals with limited dispersal ability, to colonies of microbes in biofilms, and to other genetically heterogeneous aggregates of cells, such as cancerous tumors. The spatial structure of such populations leads to greater relatedness among proximate individuals while increasing the genetic divergence between distant individuals. In this study, we analyze a Moran coa-lescent in a one-dimensional spatial model where a randomly selected individual dies and is replaced by the progeny of an adjacent neighbor in every generation. We derive a recursive system of equations using the spatial distance among haplotypes as a state variable to compute coalescent probabilities and coalescent times. The coalescent probabilities near the branch termini are smaller than in the unstructured Moran model (except for t = 1, where they are equal), corresponding to longer branch lengths and greater expected pairwise coalescent times. The lower terminal coalescent probabilities result from a spatial separation of lineages, i.e. a coalescent event between a haplotype and its neighbor in one spatial direction at time t cannot co-occur with a coalescent event with a haplotype in the opposite direction at t + 1. The concomitant increased pairwise genetic distance among randomly sampled haplotypes in spatially constrained populations could lead to incorrect inferences of recent diversifying selection or of population bottlenecks when analyzed using an unconstrained coalescent model as a null hypothesis.

genetics

Estimation of Pairwise Genetic Distances Under Independent Sampling of Segregating Sites vs. Haplotype Sampling

Genetic distance is a standard measure of variation in populations. When sequencing genomes individually, genetic distances are computed over all pairs of multilocus haplotypes in a sample. However, when next-generation sequencing methods obtain reads from heterogeneous assemblages of genomes (e.g. for microbial samples in a biofilm or cells from a tumor), individual reads are often drawn from different genomes. This means that pairwise genetic distances are calculated across independently sampled sites rather than across haplotype pairs. In this paper, we show that while the expected pairwise distance under whole haplotype sampling (WHS) is the same as with independent locus sampling (ILS), the sample variances of pairwise distance differ and depend on the direction and magnitude of linkage disequilibrium (LD) among polymorphic sites. We derive a weighted LD value that, when positive, predicts higher sample variance in estimated genetic distance for WHS. Weighted LD is positive when on average, the most common alleles at two loci are in positive LD. Using individual-based simulations of an infinite sites model under Fisher-Wright genetic drift, variances of estimated genetic distance are found to be almost always higher under WHS than under ILS, suggesting a reduction in estimation error when sites are sampled independently. We apply these results to haplotype frequencies from a lung cancer tumor to compute weighted LD and the variances in estimated genetic distance under ILS vs. WHS, and find that the the relative magnitudes of variances under WHS vs. ILS are sensitive to sampled allele frequencies.

genetics