bioRxiv · 10.1101/2023.05.21.541623
Which Variable Should Be Dependent in Phylogenetic Generalized Least Squares Regression Analysis
Abstract
Phylogenetic generalized least squares (PGLS) regression is widely used to examine evolutionary associations while accounting for phylogenetic non-independence, but it requires designating one trait as the dependent variable and the other as the independent variable. When causal relationships between traits are unclear, choosing which trait serves as the dependent variable becomes a practical concern. While studying the evolutionary relationship between bacterial growth rate and CRISPR-Cas content, we noticed that switching the roles of dependent and independent variables in PGLS analyses could lead to inconsistent conclusions in a substantial proportion of cases. To confirm this observation, we conducted 16,000 simulations of trait evolution along binary trees with 100 terminal nodes under different evolutionary models. PGLS regressions using Pagels {lambda} model were applied to each simulation, and the results consistently showed that swapping the dependent and independent variables can lead to inconsistent outcomes. We evaluated seven potential criteria for selecting the dependent variable, including log-likelihood, AIC, R2, p-value, Pagels {lambda}, Blombergs K, and the estimated {lambda} in the Pagels {lambda} model. Among these, Pagels {lambda}, Blombergs K, and the estimated {lambda} performed equally well and outperformed the others in selecting the dependent variable, providing a reliable basis for PGLS analyses when causal direction between traits is unclear.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chen, Z.-L., Guo, H.-J., Niu, D.-K.. 2023-05-23. Which Variable Should Be Dependent in Phylogenetic Generalized Least Squares Regression Analysis. https://doi.org/10.1101/2023.05.21.541623
Cite the original work for its findings. Save a collection to share your selection of sources.