Search bioRxivSearch

Biology subjects

Gascuel, O.

Publications and source records attributed to Gascuel, O..

4 recordsLinked to original sources

A Fast Likelihood Method to Reconstruct and Visualize Ancestral Scenarios

The reconstruction of ancestral scenarios is widely used to study the evolution of characters along a phylogenetic tree. In the likelihood framework one commonly uses the marginal posterior probabilities of the character states, and the joint reconstruction of the most likely scenario. Both approaches are somewhat unsatisfactory. Marginal reconstructions provide users with state probabilities, but these are difficult to interpret and visualize, while joint reconstructions select a unique state for every tree node and thus do not reflect the uncertainty of inferences.\n\nWe propose a simple and fast approach, which is in between these two extremes. We use decision-theory concepts and the Brier criterion to associate each node in the tree to a set of likely states. A unique state is predicted in the tree regions with low uncertainty, while several states are predicted in the uncertain regions, typically around the tree root. To visualize the results, we cluster the neighboring nodes associated to the same states and use graph visualization tools. The method is implemented in the PastML program and web server.\n\nThe results on simulated data consistently show the accuracy and robustness of the approach. The method is applied to large tree comprising 3,619 sequences from HIV-1M subtype C sampled worldwide, which is processed in a few minutes. Results are very convincing: we retrieve and visualize the main transmission routes of HIV-1C; we demonstrate that drug resistance mutations mostly emerge independently under treatment pressure, but some resistance clusters are found, corresponding to transmissions among untreated patients.

evolutionary biology

Distribution and asymptotic behavior of the phylogenetic transfer distance

The transfer distance (TD) was introduced in the classification framework and studied in the context of phylogenetic tree matching. Recently, Lemoine et al. (2018) showed that TD can be a powerful tool to assess the branch support of phylogenies with large data sets, thus providing a relevant alternative to Felsensteins bootstrap. This distance allows a reference branch {beta} in a reference tree [T] to be compared to a branch b from another tree T, both on the same set of n taxa. The TD between these branches is the number of taxa that must be transferred from one side of b to the other in order to obtain {beta}. By taking the minimum TD from {beta} to all branches in T we define the transfer index, denoted by{phi} ({beta}, T), measuring the degree of agreement of {beta} with T. Let us consider a reference branch {beta} having p tips on its light side and define the transfer support (TS) as 1 -{phi} ({beta}, T)/(p - 1). The aim of this article is to provide evidence that p 1 is a meaningful normalization constant in the definition of TS, and measure the statistical significance of TS, assuming that {beta} is compared to a tree T drawn according to a null model. We obtain several results that shed light on these questions in a number of settings. In particular, we study the asymptotic behavior of TS when n tends to {infty}, and fully characterize the distribution of{phi} when T is a caterpillar tree.

evolutionary biology

Boosting Felsenstein Phylogenetic Bootstrap

Felsensteins article describing the application of the bootstrap to evolutionary trees, is one of the most cited papers of all time. That statistical method, based on resampling and replications, is used extensively to assess the robustness of phylogenetic inferences. However, increasing numbers of sequences are now available for a wide variety of species, and phylogenies with hundreds or thousands of taxa are becoming routine. In that framework, Felsensteins bootstrap tends to yield very low supports, especially on deep branches. We propose a revised version, in which the presence of inferred branches in replications is measured using a gradual \"transfer\" distance, as opposed to the original version using a binary presence/absence index. The resulting supports are higher, while not inducing falsely supported branches. Our method is applied to large simulation, mammal and HIV datasets, for which it reveals the phylogenetic signal, while Felsensteins bootstrap fails to do so.

evolutionary biology

Improving pairwise comparison of protein sequences with domain co-occurrence

MotivationComparing and aligning protein sequences is an essential task in bioinformatics. More specifically, local alignment tools like BLAST are widely used for identifying conserved protein sub-sequences, which likely correspond to protein domains or functional motifs. However, to limit the number of false positives, these tools are used with stringent sequence-similarity thresholds and hence can miss several hits, especially for species that are phylogenetically distant from reference organisms. A solution to this problem is then to integrate additional contextual information to the procedure.\n\nResultsHere, we propose to use domain co-occurrence to increase the sensitivity of pairwise sequence comparisons. Domain co-occurrence is a strong feature of proteins, since most protein domains tend to appear with a limited number of other domains on the same protein. We propose a method to take this information into account in a typical BLAST analysis and to construct new domain families on the basis of these results. We used Plasmodium falciparum as a case study to evaluate our method. The experimental findings showed an increase of 16% of the number of significant BLAST hits and an increase of 28% of the proteome area that can be covered with a domain. Our method identified 2473 new domains for which, in most cases, no model of the Pfam database could be linked. Moreover, our study of the quality of the new domains in terms of alignment and physicochemical properties show that they are close to that of standard Pfam domains.\n\nAvailabilitySoftware implementing the proposed approach and the Supplementary Data are available at: https://gite.lirmm.fr/menichelli/pairwise-comparison-with-cooccurrence

bioinformatics