Search bioRxiv⌕ Search

Biology subjects

Fosse, S.

Publications and source records attributed to Fosse, S..

2 recordsLinked to original sources

Benchmarking BEAGLE to find optimal parameters for BEAST X

Bayesian phylogenetic analyses are notoriously time-consuming, largely because exploring the posterior distribution requires computing Felsensteins likelihood. The BEAGLE library is a high-performance computational tool that dramatically accelerates the calculation of such likelihoods by leveraging parallel processing on GPUs, multicore CPUs, and SSE vectorisation. Here we present results from benchmarking a widely popular phylogenetics package, BEAST X, using BEAGLE integration, focusing on how hardware allocation affects running times. We demonstrate substantial differences among BEAGLE settings on real Dengue Virus (DENV) data, both with and without partitioning. Using simulated sequences, we establish guidelines for GPU usage in BEAST X runs. These guidelines can be used for effective resource allocation for empirical analyses and simulation studies.

bioinformatics↗

The limits of Bayesian estimates of divergence times in measurably evolving populations

Bayesian inference of divergence times for extant species using molecular data is an unconventional statistical problem: Divergence times and molecular rates are confounded, and only their product, the molecular branch length, is statistically identifiable. This means we must use priors on times and rates to break the identifiability problem. As a consequence, there is a lower bound in the uncertainty that can be attained under infinite data for estimates of evolutionary timescales using the molecular clock. With infinite data (i.e., an infinite number of sites and loci in the alignment) uncertainty in ages of nodes in phylogenies increases proportionally with their mean age, such that older nodes have higher uncertainty than younger nodes. On the other hand, if extinct taxa are present in the phylogeny, and if their sampling times are known (i.e., heterochronous data), then times and rates are identifiable and uncertainties of inferred times and rates go to zero with infinite data. However, in real heterochronous datasets (such as viruses and bacteria), alignments tend to be small and how much uncertainty is present and how it can be reduced as a function of data size are questions that have not been explored. This is clearly important for our understanding of the tempo and mode of microbial evolution using the molecular clock. Here we conducted extensive simulation experiments and analyses of empirical data to develop the infinite-sites theory for heterochronous data. Contrary to expectations, we find that uncertainty in ages of internal nodes scales positively with the distance to their closest tip with known age (i.e., calibration age), not their absolute age. Our results also demonstrate that estimation uncertainty decreases with calibration age more slowly in datasets with more, rather than fewer site patterns, although overall uncertainty is lower in the former. Our statistical framework establishes the minimum uncertainty that can be attained with perfect calibrations and sequence data that are effectively infinitely informative. Finally, we discuss the implications for viral sequence datasets. In a vast majority of cases viral data from outbreaks is not sufficiently informative to display infinite-sites behaviour and thus all estimates of evolutionary timescales will be associated with a degree of uncertainty that will depend on the size of the dataset, its information content, and the complexity of the model. We anticipate that our framework is useful to determine such theoretical limits in empirical analyses of microbial outbreaks. Significance statementGenetic sequences are routinely used to date the origins of outbreaks and the divergence of viral and bacterial lineages, but every such date carries uncertainty, and it has been unclear how much of that uncertainty more sequence data can remove. We show that for pathogens sampled repeatedly over time, precision depends not on how old an event is but on how close it lies to the nearest sample of known date, and that even large influenza and hepatitis B virus datasets fall far short of the precision that is theoretically attainable. This provides a simple diagnostic for judging how much confidence a molecular date deserves, and whether sequencing more genomes would improve its precision.

bioinformatics↗