Search bioRxivSearch

Biology subjects

Hejase, H.

Publications and source records attributed to Hejase, H..

2 recordsLinked to original sources

Non-parametric and semi-parametric support estimation using SEquential RESampling random walks on biomolecular sequences

Non-parametric and semi-parametric resampling procedures are widely used to perform support estimation in computational biology and bioinformatics. Among the most widely used methods in this class is the standard bootstrap method, which consists of random sampling with replacement. While not requiring assumptions about any particular parametric model for resampling purposes, the bootstrap and related techniques assume that sites are independent and identically distributed (i.i.d.). The i.i.d. assumption can be an over-simplification for many problems in computational biology and bioinformatics. In particular, sequential dependence within biomolecular sequences is often an essential biological feature due to biochemical function, evolutionary processes such as recombination, and other factors.\n\nTo relax the simplifying i.i.d. assumption, we propose a new non-parametric/semi-parametric sequential resampling technique that generalizes \"Heads-or-Tails\" mirrored inputs, a simple but clever technique due to Landan and Graur. The generalized procedure takes the form of random walks along either aligned or unaligned biomolecular sequences. We refer to our new method as the SERES (or \"SEquential RESampling\") method.\n\nTo demonstrate the flexibility of the new technique, we apply SERES to two different applications - one involving aligned inputs and the other involving unaligned inputs. Using simulated and empirical data, we show that SERES-based support estimation yields comparable or typically better performance compared to state-of-the-art methods for both applications.

bioinformatics

FastNet: Fast and accurate inference of phylogenetic networks using large-scale genomic sequence data

An emerging discovery in phylogenomics is that interspecific gene flow has played a major role in the evolution of many different organisms. To what extent is the Tree of Life not truly a tree reflecting strict \"vertical\" divergence, but rather a more general graph structure known as a phylogenetic network which also captures \"horizontal\"gene flow? The answer to this fundamental question not only depends upon densely sampled and divergent genomic sequence data, but also compu-tational methods which are capable of accurately and efficiently inferring phylogenetic networks from large-scale genomic sequence datasets. Re-cent methodological advances have attempted to address this gap. How-ever, in the 2016 performance study of Hejase and Liu, state-of-the-art methods fell well short of the scalability requirements of existing phy-logenomic studies.\n\nThe methodological gap remains: how can phylogenetic networks be ac-curately and efficiently inferred using genomic sequence data involving many dozens or hundreds of taxa? In this study, we address this gap by proposing a new phylogenetic divide-and-conquer method which we call FastNet. We conduct a performance study involving a range of evolu-tionary scenarios, and we demonstrate that FastNet outperforms state-of-the-art methods in terms of computational efficiency and topological accuracy.

bioinformatics