Search bioRxiv⌕ Search

Biology subjects

Wiens, B. J.

Publications and source records attributed to Wiens, B. J..

2 recordsLinked to original sources

That's not a Hybrid: How to Distinguish Patterns of Admixture and Isolation-by-Distance

Describing naturally occurring genetic variation is a fundamental goal of molecular phylogeography and population genetics. Popular methods for this task include STRUCTURE, a model-based algorithm that assigns individuals to genetic clusters, and Principal Component Analysis (PCA), a parameter-free method. The ability of STRUCTURE to infer mixed ancestry makes it popular for documenting natural hybridization, which is of considerable interest to evolutionary biologists, given that such systems provide a window into the speciation process. Yet STRUCTURE can produce misleading results when its underlying assumptions are violated, like when genetic variation is distributed continuously. To test the ability of STRUCTURE and PCA to accurately distinguish admixture from continuous variation, we use forward-time simulations to generate population genetic data under three demographic scenarios: two involving admixture and one with isolation-by-distance (IBD). STRUCTURE and PCA alone cannot distinguish admixture from IBD, but complementing these analyses with triangle plots, which visualize hybrid index against interclass heterozygosity, provides more accurate inference of demographic history. We demonstrate that triangle plots are robust to missing data, while STRUCTURE and PCA are not, and show that setting a low allele frequency difference threshold for AIM identification can accurately characterize the relationship between hybrid index and interclass heterozygosity across demographic histories of admixture and range expansion. While STRUCTURE and PCA provide useful summaries of genetic variation, results should be paired with triangle plots before admixture is inferred.

evolutionary biology↗

triangulaR: an R package for identifying AIMs and building triangle plots using SNP data from hybrid zones

Hybridization provides a window into the speciation process and reshuffles parental alleles to produce novel recombinant genotypes. The presence or absence of specific hybrid classes across a hybrid zone can provide support for various modes of reproductive isolation. Early generation hybrid classes can be distinguished by their combination of hybrid index and interclass heterozygosity, which can be estimated with molecular data. Hybrid index and interclass heterozygosity are routinely calculated for studies of hybrid zones, but available resources for next-generation sequencing datasets are computationally demanding and tools for visualizing those metrics as a triangle plot are lacking. Here, we provide a resource for identifying ancestry- informative markers (AIMs) from SNP datasets, calculating hybrid index and interclass heterozygosity, and visualizing the relationship as a triangle plot. Our methods are implemented in the R package triangulaR. We validate our methods by simulating genetic data for a hybrid zone between parental groups at low, medium, and high levels of divergence. We find that accurate and precise estimates of hybrid index and interclass heterozygosity can be obtained with sample sizes as low as five individuals per parental group. We explore various allele frequency difference thresholds for AIM identification, and how this threshold influences the accuracy and precision of hybrid index and interclass heterozygosity estimates. We contextualize interpretation of triangle plots by describing the theoretical expectations for covariance of hybrid index and interclass heterozygosity under Hardy-Weinberg Equilibrium and provide recommendations for best practices for identifying AIMs and building triangle plots.

evolutionary biology↗