Search bioRxivSearch

Biology subjects

Vineet Bafna

Publications and source records attributed to Vineet Bafna.

2 recordsLinked to original sources

CLEAR: Composition of Likelihoods for Evolve And Resequence Experiments

Experimental evolution (EE) studies are powerful tools for observing molecular evolution \"in-action\" from populations sampled in controlled and natural environments. The advent of next generation sequencing technologies has made whole-genome and whole-population sampling possible, even for eukaryotic organisms with large genomes, and allowed us to locate the genes and variants responsible for genetic adaptation. While many computational tests have been developed for detecting regions under selection, they are mainly designed for static (single time) data, and work best when the favored allele is close to fixation.\n\nEE studies provide samples over multiple time points, underscoring the need for tools that can exploit the data. At the same time, EE studies are constrained by the limited time span since onset of selection, depending upon the generation time for the organism. This constraint impedes adaptation studies, as the population can only be evolve-and-resequenced for a small number of generations relative to the fixation time of the favored allele. Moreover, coverage in pool-sequenced experiments varies across replicates and time points, leading to heterogeneous ascertainment bias in measuring population allele frequency across different measurements.\n\nIn this article, we directly address these issues while developing tools for identifying selective sweep in pool-sequenced EE of sexual organisms and propose Composition of Likelihoods for Evolve-And-Resequence experiments (CO_SCPLOWLEARC_SCPLOW). Extensive simulations show that CO_SCPLOWLEARC_SCPLOW achieves higher power in detecting and localizing selection over a wide range of parameters. In contrast to existing methods, the CO_SCPLOWLEARC_SCPLOW statistics are robust to variation of coverage. CO_SCPLOWLEARC_SCPLOW also provides robust estimates of model parameters, including selection strength and overdominance, as byproduct of the statistical test, while being orders of magnitude faster than existing methods. Finally, we apply the CO_SCPLOWLEARC_SCPLOW statistic to data from a study of D. melanogaster adaptation to alternating temperatures and discover enrichment of genes related to \"response to heat\" and \"cold acclimation\".

Evolutionary Biology

Predicting Carriers of Ongoing Selective Sweeps Without Knowledge of the Favored Allele

Methods for detecting the genomic signatures of natural selection have been heavily studied, and they have been successful in identifying many selective sweeps. For most of these sweeps, the favored allele remains unknown, making it difficult to distinguish carriers of the sweep from non-carriers. In an ongoing selective sweep, carriers of the favored allele are likely to contain a future most recent common ancestor. Therefore, identifying them may prove useful in predicting the evolutionary trajectory -- for example, in contexts involving drug-resistant pathogen strains or cancer subclones. The main contribution of this paper is the development and analysis of a new statistic, the Haplotype Allele Frequency (HAF) score. The HAF score, assigned to individual haplotypes in a sample, naturally captures many of the properties shared by haplotypes carrying a favored allele. We provide a theoretical framework for computing expected HAF scores under different evolutionary scenarios, and we validate the theoretical predictions with simulations. As an application of HAF score computations, we develop an algorithm (PreCIOSS: Predicting Carriers of Ongoing Selective Sweeps) to identify carriers of the favored allele in selective sweeps, and we demonstrate its power on simulations of both hard and soft sweeps, as well as on data from well-known sweeps in human populations.\n\nAuthor summaryMethods for detecting the genomic signatures of natural selection have been heavily studied, and they have been successful in identifying genomic regions under positive selection. However, methods that detect positive selective sweeps do not typically identify the favored allele, or even the haplotypes carrying the favored allele. The main contribution of this paper is the development and analysis of a new statistic (the HAF score), assigned to individual haplotypes. Using both theoretical analyses and simulations, we describe how the HAF scores differ for carriers and non-carriers of the favored allele, and how they change dynamically during a selective sweep. We also develop an algorithm, PreCIOSS, for separating carriers and non-carriers. Our tool has broad applicability as carriers of the favored allele are likely to contain a future most recent common ancestor. Therefore, identifying them may prove useful in predicting the evolutionary trajectory -- for example, in contexts involving drug-resistant pathogen strains or cancer subclones.

Evolutionary Biology