bioRxiv · 10.1101/2024.03.25.586640
Accelerated Bayesian inference of population size history from recombining sequence data
Abstract
I present O_SCPLOWPHLASHC_SCPLOW, a new Bayesian method for inferring population history from whole genome sequence data. O_SCPLOWPHLASHC_SCPLOW is population history learning by averaging sampled histories: it works by drawing random, low-dimensional projections of the coalescent intensity function from the posterior distribution of a O_SCPLOWPSMCC_SCPLOW-like model, and averaging them together to form an accurate and adaptive size history estimator. On simulated data, O_SCPLOWPHLASHC_SCPLOW tends to be faster and have lower error than several competing methods including O_SCPLOWSMCC_SCPLOW++, O_SCPLOWMSMCC_SCPLOW2, and FO_SCPLOWITC_SCPLOWCO_SCPLOWOALC_SCPLOW. Moreover, it provides a full posterior distribution over population size history, leading to automatic uncertainty quantification of the point estimates, as well to new Bayesian testing procedures for detecting population structure and ancient bottlenecks. On the technical side, the key advance is a novel algorithm for computing the score function (gradient of the log-likelihood) of a coalescent hidden Markov model: when there are M hidden states, the algorithm requires. [O](M 2) time and. [O](1) memory per decoded position, the same cost as evaluating the log-likelihood itself using the naive forward algorithm. This algorithm is combined with a hand-tuned implementation that fully leverages the power of modern GPU hardware, and the entire method has been released as an easy-to-use Python software package.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Terhorst, J.. 2024-03-27. Accelerated Bayesian inference of population size history from recombining sequence data. https://doi.org/10.1101/2024.03.25.586640
Cite the original work for its findings. Save a collection to share your selection of sources.