Search bioRxiv⌕ Search

Biology subjects

Temple, S. D.

Publications and source records attributed to Temple, S. D..

2 recordsLinked to original sources

Identity-by-descent in large samples

If two haplotypes share the same alleles for an extended gene tract, these haplotypes are likely to be derived identical-by-descent from a recent common ancestor. Identity-by-descent segment lengths are correlated via unobserved ancestral tree and recombination processes, which commonly presents challenges to the derivation of theoretical results in population genetics. We show that the proportion of detectable identity-by-descent segments around a locus is normally distributed when the sample size and the scaled population size are large. We generalize this central limit theorem to cover flexible demographic scenarios, multi-way identity-by-descent segments, and multivariate identity-by-descent rates. We use efficient simulations to study the distributional behavior of the detectable identity-by-descent rate. One consequence of non-normality in finite samples is that a genome-wide scan looking for excess identity-by-descent rates may be subject to anti-conservative control of family-wise error rates. HighlightsO_LIWe show the asymptotic normality of the detectable identity-by-descent rate, a mean of correlated binary random variables that arises in population genetics studies. C_LIO_LIWe generalize our main central limit theorem to cover scenarios of nonconstant population sizes, multi-way identity-by-descent segments, and identity-by-descent rates of multiple samples from the same population. C_LIO_LIIn enormous simulation studies, we use an efficient algorithm to characterize distributional properties of the detectable identity-by-descent rate. C_LI

genetics↗

Modeling recent positive selection in Americans of European ancestry

Recent positive selection can result in an excess of long identity-by-descent (IBD) haplotype segments. The statistical methods that we propose here address three major objectives in studying selective sweeps: scanning for regions of interest, identifying possible sweeping alleles, and estimating a selection coefficient s. First, we implement a selection scan to locate regions of excess IBD rate. Second, we develop a statistic to rank alleles that are in strong linkage disequilibrium with a putative sweeping allele. We aggregate these scores to estimate the allele frequency of the sweeping allele, even if it is not genotyped. Third, we propose an estimator for the selection coefficient and quantify uncertainty using the parametric bootstrap. Comparing against state-of-the-art methods in extensive simulations, we show that our methods are better at identifying sweeping alleles that are at low frequency and at estimating S when S [≥] 0.015. We apply these methods to study positive selection in European ancestry samples from the TOPMed project. We analyze eight loci where the IBD rate is more than four standard deviations above the population median. The IBD rate at LCT is thirty-five standard deviations above the population median, and our estimates of its selection coefficient imply strong selection within the past two hundred generations. Overall, we present robust and accurate approaches to study very recent adaptive evolution without knowing the identity of the causal allele or using time series data.

genetics↗