Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.07.15.738685

Phylogenize2: robust phylogenetic methods link genes to phenotypes across host-associated and environmental microbiomes

Abstract

In microbiome studies, associations between microbial functions and the environment are often confounded by phylogeny. While some methods explicitly account for this confounder, they require information about genome content, limiting their use in biomes where few genomes have been available. To make these methods more universally accessible, we have developed Phylogenize2, a redesigned phylogeny-aware tool for linking microbial gene families to abundance phenotypes. Phylogenize2 integrates large metagenome-assembled genome collections, including both biome-specific collections from MGnify and a broadly sampled general purpose database, GlobDB, to substantially expand species coverage, allowing its application in environments like the mouse gut and ocean. In addition, by default, Phylogenize2 uses a new robust phylogenetic testing framework that has been optimized for microbial abundance data, while also allowing the use of other comparative methods such as POMS. In an experimental mouse study, Phylogenize2 identifies that Muribaculaceae with higher abundance on a high-fat diet are enriched for proteins in the thioredoxin family, with likely roles in oxidative stress. When we apply Phylogenize2 to a polar ocean study, we find that a molybdenum-dependent PaoABC/YagTSR-like aldehyde oxidoreductase system differentiates mesopelagic from surface-dwelling Flavobacteriaceae, suggesting that aldehyde detoxification may be important for organisms that degrade marine snow. Together, these results show that Phylogenize2 expands phylogeny-aware microbiome analysis beyond the human gut and can provide insight into the genetic basis of microbiome-encoded traits in diverse environments. ImportanceMicrobiome studies often set out to identify which microbes are more or less abundant across environments, but these patterns can be difficult to interpret. Phylogenize2 is an open-source software package that allows researchers to ask whether individual microbial gene families are associated with the environment across independent branches of the microbial tree of life. By incorporating large collections of genomes from uncultivated microbes, as well as modern statistical methods designed for microbial abundance data, Phylogenize2 makes this approach practical for microbiomes beyond the human gut, including in model organisms like lab mice and free-living environments like the ocean. We also provide a pipeline that allows the use of new genome collections. In two case studies, we demonstrate that Phylogenize2 effectively prioritizes specific genes and pathways from metagenomic data, thereby leading researchers from changes in microbial abundance to more biologically interpretable explanations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Kananen, K., Tran, N., Bradley, P. H.. 2026-07-16. Phylogenize2: robust phylogenetic methods link genes to phenotypes across host-associated and environmental microbiomes. https://doi.org/10.64898/2026.07.15.738685

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Physical priors improve performance of structure-based binding affinity models

Structure-based drug discovery is a widely used paradigm for the rational design of novel small molecule therapeutics. However, the benefits conferred by the use of structural information has seen limited adoption in machine learning, where ligand-only ("2D") models are still the industry standard for molecular property or binding affinity prediction. Structure-based ("3D") ML models for binding-affinity prediction promise to present a clear advantage, but have not yet overtaken existing 2D models. Here, we show that physics-based priors can improve predictive performance of structure-based models by comparing different model architectures with varying physical priors on several prediction tasks. We present the Modular Training and Evaluation of Neural Networks (mtenn) package, where we decompose affinity prediction into separate steps of embedding structure into learned representations and combining those embeddings into a predicted binding affinity. We consider both E(3)-invariant and E(3)-equivariant architectures to determine the importance of encoding roto-translational inductive biases, as well as different methods for combining learned embeddings. By first optimizing several aspects of model construction using the general purpose PDBBind dataset, we are able to improve the performance and data efficiency of structure-based models. When subsequently trained and evaluated on the COVID Moonshot small molecule drug discovery dataset, our tuned models perform on par with industry standard ligand-only models. Our decomposed model framework highlights that encoding some physical priors improves model performance, while more complex biases such as equivariance offer limited benefit. Additionally, structure-based models generalize better to an unseen target and display higher training efficiency. Overall, these results emphasize that structure-based models benefit from their ability to incorporate physics-informed constraints, giving promising directions for model architecture development. These results also suggest that the strength of these models may be in tasks specifically aimed at generalizability, providing guidelines for their use in early-stage drug discovery campaigns.

bioinformatics↗

Rapid Shift Toward Pulsed Field Ablation and Precision Risk Stratification in High-Impact Atrial Fibrillation Research

Conventional bibliometrics rely on lifetime citations, obscuring immediate shifts in cardiovascular research paradigms. To track emerging trends in atrial fibrillation management, we performed a comparative bibliometric analysis of the 50 highest-cited original research articles per year in OpenAlex topic T10065 across consecutive 2023 (Class of 2025) and 2024 (Class of 2026) publication cohorts. Articles were ranked using a fixed 18-month post-publication citation window, and extracted concepts were normalized into 719 canonical topics and 32 parent themes using large language model curation. Concept frequency tracking demonstrated a swift technological shift, with pulsed field ablation showing the largest topic frequency increase (+0.08, from 0.30 to 0.38) to become the leading canonical topic in 2024, displacing conventional thermal pulmonary vein isolation (-0.18, 0.40 to 0.22). Simultaneously, stroke prevention focus shifted toward refined predictive modeling, with increases in Risk Stratification and Predictive Models (+0.08, 0.30 to 0.38) and CHA2DS2-VASc scoring (+0.08, from 0.10 to 0.18). High-impact atrial fibrillation research is rapidly pivoting toward non-thermal ablation safety profiling and precision risk stratification, highlighting the utility of fixed-window concept mining for capturing real-time scientific evolution. Online explorer of the result is available at https://pri.pepkio.com.

bioinformatics↗

Integrating Genomic Annotations and Traits Dependencies for single-nucleotide polymorphisms Prioritization with Causal Concept Bottleneck Models

Predicting common traits from single-nucleotide polymorphism (SNPs) data is challenging due to polygenicity, small effect sizes, and the presence of potentially mediated or spurious cross-trait associations. We propose a modeling approach that combines genomic annotations with known cross-trait relations by leveraging Causally Reliable Concept Bottleneck Models (C2BM), a deep learning architecture that factors the joint trait distribution over a graph of interpretable concepts. This design allows trait predictions to leverage information from other observed traits in addition to genomic inputs. Furthermore, the interpretable architecture of the model enables us to investigate how specific trait-trait relationships influence SNP-level predictions. We evaluate the approach on a multi-trait GWAS dataset covering five traits and show that C2BM improves predictions when ground-truth labels for related traits are available. Moreover, by analyzing variations in how trait-trait relationships influence predictions, we postulate that such differences may reflect the presence or absence of shared genetic mechanisms or indirect effects. Accepted at the CIBB 2026 conference (https://cibb2026.teralab.ai/)

bioinformatics↗