Search bioRxivSearch

Biology subjects

Stadler, T.

Publications and source records attributed to Stadler, T..

15 recordsLinked to original sources

A Multi-State Birth-Death model for Bayesian inference of lineage-specific birth and death rates

Heterogeneous populations can lead to important differences in birth and death rates across a phylogeny Taking this heterogeneity into account is thus critical to obtain accurate estimates of the underlying population dynamics. We present a new multi-state birth-death model (MSBD) that can estimate lineage-specific birth and death rates. For species phylogenies, this corresponds to estimating lineage-dependent speciation and extinction rates. Contrary to existing models, we do not require a prior hypothesis on a trait driving the rate differences and we allow the same rates to be present in different parts of the phylogeny. Using simulated datasets, we show that the MSBD model can reliably infer the presence of multiple evolutionary regimes, their positions in the tree, and the birth and death rates associated with each. We also present a re-analysis of two empirical datasets and compare the results obtained by MSBD and by the existing software BAMM. The MSBD model is implemented as a package in the Bayesian inference software BEAST2, which allows joint inference of the phylogeny and the model parameters.\n\nSignificance statementPhylogenetic trees can inform about the underlying speciation and extinction processes within a species clade. Many different factors, for instance environmental changes or morphological changes, can lead to differences in macroevolutionary dynamics within a clade. We present here a new multi-state birth-death (MSBD) model that can detect these differences and estimate both the position of changes in the tree and the associated macroevolutionary parameters. The MSBD model does not require a prior hypothesis on which trait is driving the changes in dynamics and is thus applicable to a wide range of datasets. It is implemented as an extension to the existing framework BEAST2.

evolutionary biology

Inference of species histories in the presence of gene flow

When populations become isolated, members of these populations can diverge genetically over time. This leads to genetic differences between these populations that increase over time if the isolation persists. This process can be counteracted by gene flow, i.e. when genes are exchanged between populations. In order to study the speciation processes when gene flow is present, isolation-with-migration methods have been developed. These methods typically assume that the ranked topology of the species history is already known. However, this is often not the case and the species tree is therefore of interest itself. For the inference of species trees, it is in turn often necessary to assume that there is no gene flow between co-existing species. This assumption, however, can lead to wrongly inferred speciation times and species tree topologies. We here introduce a new method that allows inference of the species tree while explicitly modelling the flow of genes between coexisting species. By using Markov chain Monte Carlo sampling, we co-infer the species tree alongside evolutionary parameters of interest. By using simulations, we show that our newly introduced approach is able to reliably infer the species trees and parameters of the isolation-with-migration model from genetic sequence data. We then use this approach to infer the species history of the mosquitoes from the Anopheles gambiae species complex. Accounting for gene flow when inferring the species history suggests a slightly different speciation order and gene flow than previously suggested.

evolutionary biology

Inferring time-dependent migration and coalescence patterns from genetic sequence and predictor data in structured populations

Population dynamics can be inferred from genetic sequence data using phylodynamic methods. These methods typically quantify the dynamics in unstructured populations or assume the parameters describing the dynamics to be constant through time in structured populations. Inference methods allowing for structured populations and parameters to vary through time involve many parameters which have to be inferred. Each of these parameters might be however only weakly informed by data. Here we introduce an approach that uses so-called predictors, such as geographic distance between locations, within a generalized linear model to inform the population dynamic parameters, namely the time-varying migration rates and effective population sizes under the marginal approximation of the structured coalescent. By using simulations, we show that we are able to reliably infer the parameters from phylogenetic trees. We then apply this framework to a previously described Ebola virus dataset. We infer incidence to be the strongest predictor for effective population size and geographic distance the strongest predictor for migration. This allows us to show not only on simulated data, but also on real data, that we are able to identify reasonable predictors. Overall, we provide a novel method that allows to identify predictors for migration rates and effective population sizes and to use these predictors to quantify migration rates and effective population sizes. Its implementation as part of the BEAST2 software package MASCOT allows to jointly infer population dynamics within structured populations, the phylogenetic tree, and evolutionary parameters.

evolutionary biology

The relationship between transmission time and clustering methods in Mycobacterium tuberculosis epidemiology

BackgroundTracking recent transmission is a vital part of controlling widespread pathogens such as Mycobacterium tuberculosis. Multiple methods with specific performance characteristics exist for detecting recent transmission chains, usually by clustering strains based on genotype similarities. With such a large variety of methods available, informed selection of an appropriate approach for determining transmissions within a given setting/time period is difficult.\n\nMethodsThis study combines whole genome sequence (WGS) data derived from 324 isolates collected 2005-2010 in Kinshasa, Democratic Republic of Congo (DRC), a high endemic setting, with phylodynamics to unveil the timing of transmission events posited by a variety of standard genotyping methods. Clustering data based on Spoligotyping, 24-loci MIRU-VNTR typing, WGS based SNP (Single Nucleotide Polymorphism) and core genome multi locus sequence typing (cgMLST) typing were evaluated.\n\nFindingsOur results suggest that clusters based on Spoligotyping could encompass transmission events that occurred over 70 years prior to sampling while 24-loci-MIRU-VNTR often represented two or more decades of transmission. Instead, WGS based genotyping applying low SNP or cgMLST allele thresholds allows for determination of recent transmission events in timespans of up to 10 years e.g. for a 5 SNP/allele cut-off.\n\nInterpretationWith the rapid uptake of WGS methods in surveillance and outbreak tracking, the findings obtained in this study can guide the selection of appropriate clustering methods for uncovering relevant transmission chains within a given time-period. For high resolution cluster analyses, WGS-SNP and cgMLST based analyses have similar clustering/timing characteristics even for data obtained from a high incidence setting.

epidemiology

Adaptive Reduction of Male Gamete Number in a Selfing Species

The number of male gametes produced is critical for reproductive success and varies greatly between and within species1-3. Evolutionary reduction of male gamete production has been widely reported in plants as a hallmark of the selfing syndrome, as well as in humans. Such a reduction may simply represent deleterious decay4-7, but evolutionary theory predicts that breeding systems could act as a major selective force on male gamete number: while large numbers of sperm should be produced in highly promiscuous species because of male-male gamete competition1, reduced sperm numbers may be advantageous at lower outcrossing rates because of the cost of gamete production. Here we used genome-wide association study (GWAS) to show a signature of polygenic selection on pollen number in the predominantly selfing plant Arabidopsis thaliana. The top associations with pollen number were significantly more strongly enriched for signatures of selection than those for ovule number and 107 phenotypes analyzed previously, indicating polygenic selection8. Underlying the strongest association, responsible for 20% of total pollen number variation, we identified the gene REDUCED POLLEN NUMBER 1 affecting cell proliferation in the male germ line. We validated its subtle but causal allelic effects using a quantitative complementation test with CRISPR-Cas9-generated null mutants in a nonstandard wild accession. Our results support polygenic adaptation underlying reduced male gamete numbers.

evolutionary biology

Phylodynamic model adequacy using posterior predictive simulations

Rapidly evolving pathogens, such as viruses and bacteria, accumulate genetic change at a similar timescale over which their epidemiological processes occur, such that it is possible to make inferences about their infectious spread using phylogenetic time-trees. For this purpose it is necessary to choose a phylodynamic model. However, the resulting inferences are contingent on whether the model adequately describes key features of the data. Model adequacy methods allow formal rejection of a model if it cannot generate the main features of the data. We present TreeModelAdequacy (TMA), a package for the popular BEAST2 software, that allows assessing the adequacy of phylodynamic models. We illustrate its utility by analysing phylogenetic trees from two viral outbreaks of Ebola and H1N1 influenza. The main features of the Ebola data were adequately described by the coalescent exponential-growth model, whereas the H1N1 influenza data was best described by the birth-death SIR model.

bioinformatics

Inferring Species Trees Using Integrative Models of Species Evolution

Evolutionary models account for either population- or species-level processes, but usually not both. We introduce a new model, the FBD-MSC, which makes it possible for the first time to integrate both the genealogical and fossilization phenomena, by means of the multispecies coalescent (MSC) and the fossilized birth-death (FBD) processes. Using this model, we reconstruct the phylogeny representing all extant and many fossil Caninae, recovering both the relative and absolute time of speciation events. We quantify known inaccuracy issues with divergence time estimates using the popular strategy of concatenating molecular alignments, and show that the FBD-MSC solves them. Our new integrative method and empirical results advance the paradigm and practice of probabilistic total evidence analyses in evolutionary biology.

evolutionary biology

Fast Bayesian Inference of Phylogenetic Models Using Parallel Likelihood Calculation and Adaptive Metropolis Sampling

O_LIPhylogenetic comparative models (PCMs) have been used to study macroevolutionary patterns, to characterize adaptive phenotypic landscapes, to quantify rates of evolution, to measure the heritability of traits, and to test various evolutionary hypotheses. A major obstacle to applying these models has been the complexity of evaluating their likelihood function. Recent works have shown that for many PCMs, the likelihood can be obtained in time proportional to the size of the tree based on post-order tree traversal, also known as pruning. Despite this progress, inferring complex multi-trait PCMs on large trees remains a time-intensive task. Here, we study parallelizing the pruning algorithm as a generic technique for speeding-up PCM-inference.\nC_LIO_LIWe implement several parallel traversal algorithms in the form of a generic C++ library for Serial and Parallel LIneage Traversal of Trees (SPLITT). Based on SPLITT, we provide examples of parallel likelihood evaluation for several popular PCMs, ranging from a single-trait Brownian motion model to complex multi-trait Ornstein-Uhlenbeck and mixed Gaussian phylogenetic models.\nC_LIO_LIUsing the phylogenetic Ornstein-Uhlenbeck mixed model (POUMM) as a showcase, we run benchmarks on up to 24 CPU cores, reporting up to an order of magnitude parallel speed-up on simulated balanced and unbalanced trees of up to 100,000 tips with up to 16 traits. Noticing that the parallel speed-up depends on multiple factors, the SPLITT library is capable to automatically select the fastest traversal strategy for a given hardware, tree-topology, and data. Combining SPLITT likelihood calculation with adaptive Metropolis sampling on real data, we show that the time for Bayesian POUMM inference on a tree of 10,000 tips can be reduced from several days to minutes.\nC_LIO_LIWe conclude that parallel pruning effectively accelerates the likelihood calculation and, thus, the statistical inference of Gaussian phylogenetic models. For time-intensive Bayesian inferences, we recommend combining this technique with adaptive Metropolis sampling. Beyond Gaussian models, the parallel tree traversal can be applied to numerous other models, including discrete trait and birth-death population dynamics models. Currently, SPLITT supports multi-core shared memory architectures, but can be extended to distributed memory architectures as well as graphical processing units.\nC_LI

evolutionary biology

Tuberculosis outbreak investigation using phylodynamic analysis

The fast evolution of pathogenic viruses has allowed for the development of phylodynamic approaches that extract information about the epidemiological characteristics of viral genomes. Thanks to advances in whole genome sequencing, they can be applied to slowly evolving bacterial pathogens like Mycobacterium tuberculosis.\n\nIn this study, we investigate the epidemiological dynamics underlying two M. tuberculosis outbreaks using phylodynamic methods. The first outbreak occurred in the Swiss city of Bern (1993-2012) and was caused by a drug-susceptible strain belonging to the phylogenetic M. tuberculosis Lineage 4. The second outbreak was caused by a multidrug-resistant (MDR) strain of Lineage 2, imported from the Wat Tham Krabok (WTK) refugee camp in Thailand into California.\n\nThere is little temporal signal in the Bern data set and moderate temporal signal in the WTK data set. We estimate an evolutionary rate of 0.0039 per single nucleotide polymorphism (SNP) per year for Bern and 0.0024 per SNP per year for WTK. Nevertheless, due to its high sampling proportion (90%) the Bern outbreak allows robust estimation of epidemiological parameters despite the poor temporal signal. Conversely, theres much uncertainty in the epidemiological estimates concerning the WTK outbreak, which has a small sampling proportion (9%). Our results suggest that both outbreaks peaked around 1990, although the Bernese outbreak was only detected in 1993, and the WTK outbreak around 2004. Furthermore, individuals were infected for a significantly longer period (around 9 years) in the WTK outbreak than in the Bern outbreak (4-5 years).\n\nOur work highlights both the limitations and opportunities of phylodynamic analysis of outbreaks involving slowly evolving pathogens: (i) estimation of the evolutionary rate is difficult on outbreak time scales and (ii) a high sampling proportion allows quantification of the age of the outbreak based on the sampling times, and thus allows for robust estimation of epidemiological parameters.

epidemiology

MASCOT: Parameter and state inference under the marginal structured coalescent approximation

MotivationThe structured coalescent is widely applied to study demography within and migration between sub-populations from genetic sequence data. Current methods are either exact but too computationally inefficient to analyse large datasets with many states, or make strong approximations leading to severe biases in inference. We recently introduced an approximation based on weaker assumptions to the structured coalescent enabling the analysis of larger datasets with many different states. We showed that our approximation provides unbiased migration rate and population size estimates across a wide parameter range.\n\nResultsWe here extend this approach by providing a new algorithm to calculate the probability of the state of internal nodes that includes the information from the full phylogenetic tree. We show that this algorithm is able to increase the probability attributed to the true node states. Furthermore we use improved integration techniques, such that our method is now able to analyse larger datasets, including a H3N2 dataset with 433 sequences sampled from 5 different locations.\n\nAvailabilityThe here presented methods are combined into the BEAST2 package MASCOT, the Marginal Approximation of the Structured COalescenT. This package can be downloaded via the BEAUti package manager. The source code is available at https://github.com/nicfel/Mascot.git.

evolutionary biology

Directly Estimating Epidemic Curves From Genomic Data

Modern phylodynamic methods interpret an inferred phylogenetic tree as a partial transmission chain providing information about the dynamic process of transmission and removal (where removal may be due to recovery, death or behaviour change). Birth-death and coalescent processes have been introduced to model the stochastic dynamics of epidemic spread under common epidemiological models such as the SIS and SIR models, and are successfully used to infer phylogenetic trees together with transmission (birth) and removal (death) rates. These methods either integrate analytically over past incidence and prevalence to infer rate parameters, and thus cannot explicitly infer past incidence or prevalence, or allow such inference only in the coalescent limit of large population size. Here we introduce a particle filtering framework to explicitly infer prevalence and incidence trajectories along with phylogenies and epidemiological model parameters from genomic sequences and case count data in a manner consistent with the underlying birth-death model. After demonstrating the accuracy of this method on simulated data, we use it to assess the prevalence through time of the early 2014 Ebola outbreak in Sierra Leone.

epidemiology

Bayesian Inference Of Species Networks From Multilocus Sequence Data

Reticulate species evolution, such as hybridization or introgression, is relatively common in nature. In the presence of reticulation, species relationships can be captured by a rooted phylogenetic network, and orthologous gene evolution can be modeled as bifurcating gene trees embedded in the species network. We present a Bayesian approach to jointly infer species networks and gene trees from multilocus sequence data. A novel birth-hybridization process is used as the prior for the species network. We assume a multispecies network coalescent (MSNC) prior for the embedded gene trees. We verify the ability of our method to correctly sample from the posterior distribution, and thus to infer a species network, through simulations. We reanalyze a large dataset of genes from closely related spruces, and verify the previously suggested homoploid hybridization event in this clade. Our method is available within the BEAST 2 add-on SpeciesNetwork, and thus provides a general framework for Bayesian inference of reticulate evolution.

evolutionary biology

External introductions helped drive and sustain the high incidence of HIV-1 in rural KwaZulu-Natal, South Africa

Despite increasing access to antiretroviral therapy, HIV incidence in rural KwaZulu-Natal communities remains among the highest ever reported in Africa. While many epidemiological factors have been invoked to explain this high incidence, widespread human mobility and movement of viral lineages between geographic locations have implicated high rates of transmission across communities. High rates of crosscommunity transmission call into question how effective increasing local coverage of antiretroviral therapy will be at preventing new infections, especially if many new cases arise from external introductions. To help address this question, we use a new phylodynamic modeling approach to estimate both changes in epidemic dynamics through time and the relative contribution of local transmission versus external introductions to overall incidence from HIV-1 subtype C phylogenies. Our phylodynamic estimates of HIV prevalence and incidence are remarkably consistent with population-based surveillance data. Our analysis also reveals that early epidemic dynamics in this population were largely driven by a wave of external introductions. More recently, we estimate that anywhere between 20-60% of all new infections arise from external introductions from outside the local community. These results highlight the power of using phylodynamic methods to study generalized HIV epidemics and the growing need to consider larger-scale regional transmission dynamics above the level of local communities when designing and testing prevention strategies.

epidemiology

POUMM: An R-package for Bayesian Inference of Phylogenetic Heritability

AO_SCPLOWBSTRACTC_SCPLOWPhylogenetic comparative methods have been used to model trait evolution, to test selection versus neutral hypotheses, to estimate optimal trait-values, and to quantify the rate of adaptation towards these optima. Several authors have proposed algorithms calculating the likelihood for trait evolution models, such as the Ornstein-Uhlenbeck (OU) process, in time proportional to the number of tips in the tree. Combined with gradient-based optimization, these algorithms enable maximum likelihood (ML) inference within seconds, even for trees exceeding 10,000 tips. Despite its useful statistical properties, ML has been criticised for being a point estimator prone to getting stuck in local optima. As an elegant alternative, Bayesian inference explores the entire information in the data and compares it to prior knowledge but, usually, runs in much longer time, even for small trees. Here, we propose an approach to use the full potential of ML and Bayesian inference, while keeping the runtime within minutes. Our approach combines (i) a new algorithm for parallel likelihood calculation; (ii) a previously published method for adaptive Metropolis sampling. In principle, the strategy of (i) and (ii) can be applied to any likelihood calculation on a tree which proceeds in a pruning-like fashion leading to enormous speed improvements. As a showcase, we implement the phylogenetic Ornstein-Uhlenbeck mixed model (POUMM) in the form of an easy-to-use and highly configurable R-package. In addition to the above-mentioned usage of comparative methods, the POUMM allows to estimate non-heritable variance and phylogenetic heritability. Using simulations and empirical data from 487 mammal species, we show that the POUMM is far more reliable in terms of unbiased estimates and false positive rate for stabilizing selection, compared to its alternative - the non-mixed Ornstein-Uhlenbeck model, which assumes a fully heritable and perfectly measurable trait. Further, our analysis reveals that the phylogenetic mixed model (PMM), which assumes neutral evolution (Brownian motion) can be a very unstable estimator of phylogenetic heritability, even if the Brownian motion assumption is only weakly violated. Our results prove the need for a simultaneous account for selection and non-heritable variance in phylogenetic evolutionary models and challenge stabilizing selection hypotheses stated in numerous macro-evolutionary studies.

evolutionary biology

The Structured Coalescent and its Approximations

Phylogenetics can be used to elucidate the movement of genes between populations of organisms, using phylogeographic methods. This has been widely done to quantify pathogen movement between different host populations, the migration history of humans, and the geographic spread of languages or the gene flow between species using the location or state of samples alongside sequence data. Phylogenies therefore offer insights into migration processes not available from classic epidemiological or occurrence data alone. Phylogeographic methods have however several known shortcomings. In particular, one of the most widely used methods treats migration the same as mutation, and therefore does not incorporate information about population demography. This may lead to severe biases in estimated migration rates for datasets where sampling is biased across populations. The structured coalescent on the other hand allows us to coherently model the migration and coalescent process, but current implementations struggle with complex datasets due to the need to infer ancestral migration histories. Thus, approximations to the structured coalescent, which integrate over all ancestral migration histories, have been developed. However, the validity and robustness of these approximations remain unclear. We present an exact numerical solution to the structured coalescent that does not require the inference of migration histories. While this solution is computationally unfeasible for large datasets, it clarifies the assumptions of previously developed approximate methods and allows us to provide an improved approximation to the structured coalescent. We have implemented these methods in BEAST2, and we show how these methods compare under different scenarios.

evolutionary biology