Search bioRxiv⌕ Search

Biology subjects

Brusselmans, M.

Publications and source records attributed to Brusselmans, M..

2 recordsLinked to original sources

Biological causes and impacts of rugged tree landscapes in phylodynamic inference

Phylodynamic analysis has been instrumental in elucidating epidemiological and evolutionary dynamics of pathogens. Bayesian phylodynamics integrates out phylogenetic uncertainty, which is typically substantial in phylodynamic datasets due to limited genetic diversity. Phylodynamic inference does not, however, scale with modern datasets, partly due to difficulties in traversing tree space. Here, we characterize tree space and landscape in phylodynamic inference and assess its impacts on analysis difficulty and key biological estimates. By running extensive Bayesian analyses of 15 classic large phylodynamic datasets and carefully analyzing the posterior samples, we find that the posterior tree landscape is diffuse yet rugged, leading to widespread tree sampling problems that usually stem from sequences in a small part of the tree. We develop clade-specific diagnostics to show that a few sequences--including putative recombinants and recurrent mutants--frequently drive the ruggedness and sampling problems, although existing data-quality tests show limited power to detect them. The sampling problems can significantly impact phylodynamic inferences or distort major biological conclusions; the impact is usually stronger on "local" estimates (e.g., introduction history) associated with particular clades than on "global" parameters (e.g., demographic trajectory) governed by general tree shape. We evaluate existing and newly-developed MCMC diagnostics, and offer strategies for optimizing phylodynamic analysis settings and mitigating sampling problem impacts. Our findings highlight the need and directions to develop efficient traversal over rugged tree landscapes, ultimately advancing scalable and reliable phylodynamics. Significance StatementBayesian phylodynamics is central to epidemiological studies, but exploring the vast and complex tree space is computationally challenging. Phylodynamic datasets comprise many highly similar sequences, sampled through time, creating a uniquely structured landscape of optimal trees. Here, we show that phylodynamic tree landscapes are often highly rugged, with multiple peaks separated by difficult-to-cross valleys. These features lead to widespread sampling problems which are often driven by a few sequences. These problems can significantly impact phylodynamic estimates, especially those associated with particular clades, distorting biological conclusions. We develop diagnostics to identify problematic sequences and provide solutions to mitigate their impacts. We offer strategies to optimize phylodynamic analysis workflows and to develop algorithms for navigating rugged landscapes, thereby advancing infectious disease investigation.

evolutionary biology↗

HIPSTR: highest independent posterior subtree reconstruction in TreeAnnotator X

In Bayesian phylogenetic and phylodynamic studies it is common to summarise the posterior distribution of trees with a time-calibrated consensus phylogeny. While the maximum clade credibility (MCC) tree is often used for this purpose, we here show that a novel consensus tree method - the highest independent posterior subtree reconstruction, or HIPSTR - contains consistently higher supported clades over MCC. We also provide faster computational routines for estimating both consensus trees in an updated version of TreeAnnotator X, an open-source software program that summarizes the information from a sample of trees and returns many helpful statistics such as individual clade credibilities contained in the consensus tree. HIPSTR and MCC reconstructions on two Ebola virus and two SARS-CoV-2 data sets show that HIPSTR yields consensus trees that consistently contain clades with higher support compared to MCC trees. The MCC trees regularly fail to include several clades with very high posterior probability ([≥] 0.95) as well as a large number of clades with moderate to high posterior probability ([≥] 0.50), whereas HIPSTR achieves near-perfect performance in this respect. HIPSTR also exhibits favorable computational performance over MCC in TreeAnnotator X. Comparison to the recently developed CCD0-MAP algorithm yielded mixed results, and requires more in-depth exploration in follow-up studies. TreeAnnotator X - which is part of the BEAST X (v10.5.0) software package - is available at https://github.com/beast-dev/beast-mcmc/releases.

bioinformatics↗