Search bioRxiv⌕ Search

Biology subjects

Guo, H.-J.

Publications and source records attributed to Guo, H.-J..

2 recordsLinked to original sources

Enhancing Phylogenetic Analysis: Integrating Nonparametric Methods for Addressing Abrupt Evolutionary Shifts

Traditional phylogenetically aware correlation methods perform well under gradual evolutionary processes. However, abrupt evolutionary shifts--or macroevolutionary jumps, characteristic of punctuated evolution--can produce extreme phylogenetically independent contrasts (PIC), leading to inflated false positives or increased false negatives in trait correlation analyses. We introduce O(D)GC (Outlier-and Distribution-Guided Correlation), a flexible workflow that identifies outliers in PICs using a distribution-free boxplot criterion and applies Spearman correlation whenever influential outliers are detected. If no outliers are detected, Pearson correlation is used--automatically for large datasets (n [≥] 30), or guided by normality testing in smaller samples. We systematically compared PIC-O(D)GC with five widely applied phylogenetic correlation methods--PIC-Pearson, PIC-MM, PGLS (phylogenetic generalized least squares), MR-PMM (multi-response phylogenetic mixed model), and Corphylo--on 322,000 simulated datasets spanning five evolutionary scenarios (two shift settings: single-trait shifts and dual-trait co-directional jumps; and three no-shift gradual evolution settings), including both fixed-depth and randomly located shifts, tested across 11 shift or noise gradients, three tree sizes (16, 128, 256 tips), and both balanced and random topologies. Overall, PIC-O(D)GC achieved error rates comparable to--or noticeably higher than--those of PIC-MM, while yielding substantially lower error rates than most alternative methods. Under no-shift conditions, it retained power similar to other methods. Analyses of three empirical datasets likewise showed that PIC-O(D)GC and PIC-MM corrected shift-induced distortions that misled conventional methods. Moreover, PIC-O(D)GC offers a conceptually simple framework and incurs markedly lower computational cost. By design, its correlation-only output provides less mechanistic detail than regression-based approaches like PGLS. However, when paired with PIC diagnostics, this outlier-guided strategy highlights evolutionary jumps, distinguishes coupled from decoupled shifts, and--via clade partitioning or tip pruning--recovers background correlations, offering biologically informative insights into how punctuated events interact with gradual trends in trait evolution.

evolutionary biology↗

Which Variable Should Be Dependent in Phylogenetic Generalized Least Squares Regression Analysis

Phylogenetic generalized least squares (PGLS) regression is widely used to examine evolutionary associations while accounting for phylogenetic non-independence, but it requires designating one trait as the dependent variable and the other as the independent variable. When causal relationships between traits are unclear, choosing which trait serves as the dependent variable becomes a practical concern. While studying the evolutionary relationship between bacterial growth rate and CRISPR-Cas content, we noticed that switching the roles of dependent and independent variables in PGLS analyses could lead to inconsistent conclusions in a substantial proportion of cases. To confirm this observation, we conducted 16,000 simulations of trait evolution along binary trees with 100 terminal nodes under different evolutionary models. PGLS regressions using Pagels {lambda} model were applied to each simulation, and the results consistently showed that swapping the dependent and independent variables can lead to inconsistent outcomes. We evaluated seven potential criteria for selecting the dependent variable, including log-likelihood, AIC, R2, p-value, Pagels {lambda}, Blombergs K, and the estimated {lambda} in the Pagels {lambda} model. Among these, Pagels {lambda}, Blombergs K, and the estimated {lambda} performed equally well and outperformed the others in selecting the dependent variable, providing a reliable basis for PGLS analyses when causal direction between traits is unclear.

evolutionary biology↗