Search bioRxivSearch

Biology subjects

O'Meara, B.

Publications and source records attributed to O'Meara, B..

3 recordsLinked to original sources

Phylotastic: improving access to tree-of-life knowledge with flexible, on-the-fly delivery of trees

(1) A comprehensive phylogeny of species, i.e., a tree of life, has potential uses in a variety of contexts in research and education. This potential is limited if accessing the tree of life requires special knowledge, complex software, or long periods of training.\n\n(2) The Phylotastic project aims to use web-services technologies to lower the barrier for accessing phylogenetic knowledge, making it as easy to get a phylogeny of species as it is to get online driving directions. In prior work, we designed an open system of web services to validate and manage species names, find phylogeny resources, extract subtrees matching a users species list, calibrate them, and mash them up with images and information from online resources.\n\n(3) Here we report a publicly accessible system for on-the-fly delivery of phylogenetic knowledge, developed with user feedback on what types of functionality are considered useful by researchers and educators. The system currently consists of a web portal that implements 3 types of workflows to obtain species phylogenies (scaled by geologic time and decorated with thumbnail images); 19 underlying web services accessible via a common registry; and toolbox code in R and Python so that others can create applications that leverage these services. These resources cover most of the use-cases identified in our analysis of user needs.\n\n(4) The Phylotastic system, accessible via http://www.phylotastic.org, provides a unique resource to access the current state of phylogenetic knowledge, useful for a variety of cases in which a tree extracted quickly from online resources (as distinct from a tree custom-made from character data) is sufficient, as it is for many casual uses of trees identified here.

bioinformatics

Hidden state models improve the adequacy of state-dependent diversification approaches using empirical trees, including biogeographical models

The state-dependent speciation and extinction models (SSE) have recently been criticized due to their high rates of \"false positive\" results and many researchers have advocated avoiding SSE models in favor of other \"non-parametric\" or \"semi-parametric\" approaches. The hidden Markov modeling (HMM) approach provides a partial solution to the issues of model adequacy detected with SSE models. The inclusion of \"hidden states\" can account for rate heterogeneity observed in empirical phylogenies and allows detection of true signals of state-dependent diversification or diversification shifts independent of the trait of interest. However, the adoption of HMM into other classes of SSE models has been hampered by the interpretational challenges of what exactly a \"hidden state\" represents, which we clarify herein. We show that HMM models in combination with a model-averaging approach naturally account for hidden traits when examining the meaningful impact of a suspected \"driver\" of diversification. We also extend the HMM to the geographic state-dependent speciation and extinction (GeoSSE) model. We test the efficacy of our \"GeoHiSSE\" extension with both simulations and an empirical data set. On the whole, we show that hidden states are a general framework that can generally distinguish heterogeneous effects of diversification attributed to a focal character.

evolutionary biology

Population Genetics Based Phylogenetics Under Stabilizing Selection for an Optimal Amino Acid Sequence: A Nested Modeling Approach

We present a new phylogenetic approach SelAC (Selection on Amino acids and Codons), whose substitution rates are based on a nested model linking protein expression to population genetics. Unlike simpler codon models which assume a single substitution matrix for all sites, our model more realistically represents the evolution of protein coding DNA under the assumption of consistent, stabilizing selection using cost-benefit approach. This cost-benefit approach allows us generate a set of 20 optimal amino acid specific matrix families using just a handful of parameters and naturally links the strength of stabilizing selection to protein synthesis levels, which we can estimate. Using a yeast dataset of 100 orthologs for 6 taxa, we find SelAC fits the data much better than popular models by 104-105 AICc units. Our results indicate there is great potential for more accurate inference of phylogenetic trees and branch lengths from already existing data through the use of nested, mechanistic models. Additional parameters estimated by SelAC indicate that a large amount of non-phylogenetic, but biologically meaningful, information can be inferred from exisiting data. For example, SelAC prediction of gene specific protein synthesis rates correlates well with both empirical (r=0.33-0.48) and other theoretical predictions (r=0.45-0.64) for multiple yeast species. SelAC also provides estimates of the optimal amino acid at each site. Finally, because SelAC is a nested approach based on clearly stated biological assumptions, future modifications, such as including shifts in the optimal amino acid sequence within or across lineages, are possible.

evolutionary biology