Search bioRxivSearch

Biology subjects

Beaulieu, J.

Publications and source records attributed to Beaulieu, J..

3 recordsLinked to original sources

Hidden state models improve the adequacy of state-dependent diversification approaches using empirical trees, including biogeographical models

The state-dependent speciation and extinction models (SSE) have recently been criticized due to their high rates of \"false positive\" results and many researchers have advocated avoiding SSE models in favor of other \"non-parametric\" or \"semi-parametric\" approaches. The hidden Markov modeling (HMM) approach provides a partial solution to the issues of model adequacy detected with SSE models. The inclusion of \"hidden states\" can account for rate heterogeneity observed in empirical phylogenies and allows detection of true signals of state-dependent diversification or diversification shifts independent of the trait of interest. However, the adoption of HMM into other classes of SSE models has been hampered by the interpretational challenges of what exactly a \"hidden state\" represents, which we clarify herein. We show that HMM models in combination with a model-averaging approach naturally account for hidden traits when examining the meaningful impact of a suspected \"driver\" of diversification. We also extend the HMM to the geographic state-dependent speciation and extinction (GeoSSE) model. We test the efficacy of our \"GeoHiSSE\" extension with both simulations and an empirical data set. On the whole, we show that hidden states are a general framework that can generally distinguish heterogeneous effects of diversification attributed to a focal character.

evolutionary biology

Insights into global planktonic diatom diversity:Comparisons between phylogenetically meaningful unitsthat account for time

Metabarcoding has offered unprecedented insights into microbial diversity. In many studies, short DNA sequences are binned into consecutively higher Linnaean ranks, and ranked groups (e.g., genera) are the units of biodiversity analyses. These analyses assume that Linnaean ranks are biologically meaningful and that identically ranked groups are comparable. We used a meta-barcode dataset for marine planktonic diatoms to illustrate the limits of this approach. We found that the 20 most abundant marine planktonic diatom genera ranged in age from 4 to 134 million years, indicating the non-equivalence of genera because some had more time to diversify than others. Still, species richness was only weakly correlated with genus age, highlighting variation in rates of speciation and/or extinction. Taxonomic classifications often do not reflect phylogeny, so genus-level analyses can include phylogenetically nested genera, further confounding rank-based analyses. These results underscore the indispensable role of phylogeny in understanding patterns of microbial diversity.

evolutionary biology

Population Genetics Based Phylogenetics Under Stabilizing Selection for an Optimal Amino Acid Sequence: A Nested Modeling Approach

We present a new phylogenetic approach SelAC (Selection on Amino acids and Codons), whose substitution rates are based on a nested model linking protein expression to population genetics. Unlike simpler codon models which assume a single substitution matrix for all sites, our model more realistically represents the evolution of protein coding DNA under the assumption of consistent, stabilizing selection using cost-benefit approach. This cost-benefit approach allows us generate a set of 20 optimal amino acid specific matrix families using just a handful of parameters and naturally links the strength of stabilizing selection to protein synthesis levels, which we can estimate. Using a yeast dataset of 100 orthologs for 6 taxa, we find SelAC fits the data much better than popular models by 104-105 AICc units. Our results indicate there is great potential for more accurate inference of phylogenetic trees and branch lengths from already existing data through the use of nested, mechanistic models. Additional parameters estimated by SelAC indicate that a large amount of non-phylogenetic, but biologically meaningful, information can be inferred from exisiting data. For example, SelAC prediction of gene specific protein synthesis rates correlates well with both empirical (r=0.33-0.48) and other theoretical predictions (r=0.45-0.64) for multiple yeast species. SelAC also provides estimates of the optimal amino acid at each site. Finally, because SelAC is a nested approach based on clearly stated biological assumptions, future modifications, such as including shifts in the optimal amino acid sequence within or across lineages, are possible.

evolutionary biology