bioRxiv · 10.64898/2025.12.08.693045
Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme
Abstract
Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured, and predictive modeling from them typically relies on ad hoc feature representation choices that obscure their impact on performance and interpretation. We present ritme, an open-source Python package that jointly optimizes microbiome-specific feature representation and model selection - combined algorithm selection and hyperparameter optimization - tailored to these data. ritme systematically searches taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment together with model class and hyperparameters, using state-of-the-art optimizers that scale from a laptop to a compute cluster. Across three real-world use cases, ritme outperformed the original study pipelines by 7-29% on the primary task metric and surpassed three AutoML baselines in six of seven comparisons, while selecting substantially fewer features and exposing how feature and model choices drive performance. Open-source and modular, ritme supports reproducible, parsimonious predictive modeling, downstream biological investigation, and extension to other multi-omics modalities.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adamov, A., Mueller, C. L., Bokulich, N.. 2025-12-11. Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme. https://doi.org/10.64898/2025.12.08.693045
Cite the original work for its findings. Save a collection to share your selection of sources.