bioRxiv · 10.1101/2025.03.17.643624
randPedPCA: Rapid approximation of principal components from large pedigrees
Abstract
BackgroundPedigrees continue to be extremely important in agriculture and conservation genetics with the pedigrees of modern breeding programmes easily comprising millions of records. The structure of such pedigrees is challenging to visualise. Being directed acyclic graphs, pedigrees can be represented as matrices. Common choices are the numerator relationship matrix, A, and the adjacency matrix T. With these matrices, the structure of pedigrees can then, in principle, be visualised via principal component analysis (PCA). However, the naive PCA of matrices for large pedigrees is challenging due to computational and memory constraints. ResultsWe present the open-access R package randPedPCA for rapid pedigree PCA using sparse matrices and randomised linear algebra. Our rapid pedigree PCA builds on the fact that matrix-vector multiplications with the numerator relationship matrix can be carried out implicitly using the extremely sparse inverse Cholesky factor of the numerator relationship matrix. We demonstrate the utility of randPedPCA using several simulated datasets of pedigrees and SNPs. We then demonstrate the performance of randPedPCA by analysing the pedigree of the UK Labrador Retriever breeding population of almost 1.5 million individuals. ConclusionsThe structure of pedigrees can be efficiently and rapidly visualised using scatter plots of principal component scores. For large pedigrees, this is considerably faster than rendering plots of a pedigree graph.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Lee, H., Craddock, R. F., Gorjanc, G., Becher, H.. 2025-03-17. randPedPCA: Rapid approximation of principal components from large pedigrees. https://doi.org/10.1101/2025.03.17.643624
Cite the original work for its findings. Save a collection to share your selection of sources.