bioRxiv · 10.1101/2021.07.13.452141
Modelling complex population structure using F-statistics and Principal Component Analysis
Abstract
Human genetic diversity is shaped by our complex history. Data-driven methods such as Principal Component Analysis (PCA) are an important population genetic tool to understand this method. Here, I contrast PCA with a set of statistics motivated by trees (F-statistics). Here, I show that these two methods are closely related, and I derive explicit connections between the two approaches. I show that F-statistics have a simple geometrical interpretation in the context of PCA, and that orthogonal projections are the key concept to establish this link. I illustrate my results on two examples, one of local, and one of global human diversity. In both examples, I find that just using the first few PCs provides good population structure is sparse, and only a few components contribute to most statistics. Based on these results, I develop novel visualizations that allow for investigating specific hypotheses, checking the assumptions of more sophisticated models. My results extend F-statistics to non-discrete populations, moving towards more complete and less biased descriptions of human genetic variation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Peter, B. M.. 2021-07-13. Modelling complex population structure using F-statistics and Principal Component Analysis. https://doi.org/10.1101/2021.07.13.452141
Cite the original work for its findings. Save a collection to share your selection of sources.