Search bioRxiv⌕ Search

Biology subjects

Zaman, T. A.

Publications and source records attributed to Zaman, T. A..

3 recordsLinked to original sources

On the robustness to gene tree rooting (or lack thereof) of triplet-based species tree estimation methods

Species tree estimation is frequently based on phylogenomic approaches that use multiple genes from throughout the genome. This process becomes particularly challenging due to gene tree heterogeneity (discordance), often resulting from Incomplete Lineage Sorting (ILS). Triplet and quartet-based approaches for species tree estimation have gained substantial attention as they are provably statistically consistent in the presence of ILS. However, unlike quartet-based methods, the limitation of rooted triplet-based methods in handling unrooted gene trees has restricted their adoption in the systematics community. Furthermore, since the induced triplet distribution in a gene tree depends on the placement of the root, the accuracy of triplet-based methods depends on the accuracy of gene tree rooting. Despite progress in developing methods for rooting unrooted gene trees, greatly understudied is the choice of rooting technique and downstream effects on species tree inference under realistic model conditions. This study involves rigorous empirical testing with different gene tree rooting approaches to establish a nuanced understanding of the impact of rooting on species tree accuracy. Moreover, we aim to investigate the conditions under which triplet-based methods provide more accurate species tree estimations than the widely-used quartet-based methods such as ASTRAL.

evolutionary biology↗

GraFusionNet: Integrating Node, Edge, and Semantic Features for Enhanced Graph Representations

Understanding complex graph-structured data is a cornerstone of modern research in fields like cheminformatics and bioinformatics, where molecules and biological systems are naturally represented as graphs. However, traditional graph neural networks (GNNs) often fall short by focusing mainly on node features while overlooking the rich information encoded in edges. To bridge this gap, we present GraFusionNet, a framework designed to integrate node, edge, and molecular-level semantic features for enhanced graph classification. By employing a dual-graph autoencoder, GraFusionNet transforms edges into nodes via a line graph conversion, enabling it to capture intricate relationships within the graph structure. Additionally, the incorporation of Chem-BERT embeddings introduces semantic molecular insights, creating a comprehensive feature representation that combines structural and contextual information. Our experiments on benchmark datasets, such as Tox21 and HIV, highlight GraFusionNets superior performance in tasks like toxicity prediction, significantly surpassing traditional models. By providing a holistic approach to graph data analysis, GraFusion-Net sets a new standard in leveraging multi-dimensional features for complex predictive tasks. CCS CONCEPTSO_LIComputing methodologies [->] Neural networks. C_LI ACM Reference FormatMd Toki Tahmid, Tanjeem Azwad Zaman, and Mohammad Saifur Rahman. 2018. GraFusionNet: Integrating Node, Edge, and Semantic Features for Enhanced Graph Representations. In Proceedings of Make sure to enter the correct conference title from your rights confirmation email (Conference acronym XX). ACM, New York, NY, USA, 9 pages. https://doi.org/XXXXXXX.XXXXXXX

molecular biology↗

wQFM-TREE: highly accurate and scalable quartet-based species tree inference from gene trees

Summary methods are becoming increasingly popular for species tree estimation from multi-locus data in the presence of gene tree discordance. ASTRAL, a leading method in this class, solves the Maximum Quartet Support Species Tree problem within a constrained solution space constructed from the input gene trees. In contrast, alternative heuristics such as wQFM and wQMC operate by taking a set of weighted quartets as input and employ a divide-and-conquer strategy to construct the species tree. Recent studies showed wQFM to be more accurate than ASTRAL and wQMC, though its scalability is hindered by the computational demands of explicitly generating and weighting {Theta}(n4) quartets. Here, we introduce wQFM-TREE, a novel summary method that enhances wQFM by circumventing the need for explicit quartet generation and weighting, thereby enabling its application to large datasets. Unlike wQFM, wQFM-TREE can also handle polytomies. Extensive simulations under diverse and challenging model conditions, with hundreds or thousands of taxa and genes, consistently demonstrate that wQFM-TREE matches or improves upon the accuracy of ASTRAL. Specifically, wQFM-TREE outperformed ASTRAL in 25 of 27 model conditions analyzed in this study involving 200-1000 taxa, with statistically significant differences in 20 of these conditions. Moreover, we applied wQFM-TREE to re-analyze the green plant dataset from the One Thousand Plant Transcriptomes Initiative. Its remarkable accuracy and scalability position wQFM-TREE as a highly competitive alternative to leading methods in the field. Additionally, the algorithmic and combinatorial innovations introduced in this study will benefit various quartet-based computations, advancing the state-of-the-art in phylogenetic estimations.

evolutionary biology↗