Search bioRxiv⌕ Search

Biology subjects

Ur-Rehman, S.

Publications and source records attributed to Ur-Rehman, S..

2 recordsLinked to original sources

Learning the Language of the Microbiome with Transformers

Self-supervised pretraining has become central to biological machine learning, yet microbiome data remains comparatively underexplored in terms of both modeling approaches and evaluation frameworks. To address this gap, we present Atlas, a pretraining dataset of 539,308 microbiome datapoints from the MGnify database. Using Atlas, we train the Waypoint family of microbiome foundation models: a series of GPT-2 style causal language models ranging from 6M to 170M parameters. We also introduce Compass, a curated benchmark of eight predictive tasks spanning biome classification, drug-microbiome interactions, drug degradation, and infant gut development. Using this benchmark, we compare the performance of Waypoint models against classical baselines and the existing MGM foundation model. Our results show that pretraining leads to consistent and significant improvements in downstream task performance, that both dataset scale and tokenization strategy impact model quality, that pretraining is essential for achieving favorable scaling behavior and that representations learned during pretraining generalise between microbiome domains. Furthermore, pretrained transformer models begin to reliably outperform classical methods once training data exceeds roughly 10,000 examples - a threshold that is attainable for modern microbiome studies. Finally, we demonstrate that the Waypoint models achieve state-of-the-art performance among microbiome foundation models. Overall, our work highlights the importance of large-scale self-supervised pretraining in this domain and establishes Atlas, Compass, and the Waypoint models as valuable resources for the research community in this emerging field.

bioinformatics↗

Breaking Through Biology's Data Wall: Expanding the Known Tree of Life by Over 10x using a Global Biodiscovery Pipeline

Advancements in the life sciences have always been built upon our collective understanding of life on Earth. Now, the rise of generative biology - the use of AI foundation models to design, generate, and annotate proteins, pathways and therapeutics - is creating unprecedented demand for large, diverse biological sequence datasets. While a limited subset of such data can be generated in clinical or laboratory settings, the vast majority of the training data for unsupervised models must be sourced from the natural world - the product of nearly four billion years of evolutionary history. However, the public databases that currently supply this data, while foundational to research, were established to aggregate results from academic experiments, not as training datasets for machine learning. Their human-centric data structure limits model performance due to redundancy, taxonomic and geographic bias, limited biological context, and inconsistent provenance. With 68% of all sequence data in the SRA database coming from just 5 species, this is one of the most severe class imbalance problems ever encountered in AI. Legal and infrastructural constraints further exacerbate this bottleneck. To address these limitations and support scalable model training, we introduce BaseData: the largest and fastest-growing biological sequence database ever built, and the first purpose-built for training foundation models. As of late 2024, BaseData contained 9.8 billion novel genes, representing more than a 10-fold expansion in known protein diversity after accounting for redundancy. BaseData also contains more than 1 million species not represented in other genomic databases. Its partnership-driven data supply chain across 26 countries and autonomous regions enables growth of up to 2 billion novel genes per month, far exceeding public repositories. All data is collected under benefit sharing agreements using standardized protocols and structured using graph-based, ontology-rich metadata that preserves evolutionary context. BaseData represents a new, ethically-grounded infrastructure for training biological foundation models, complementing public efforts and enabling the next era of generative biology. Short AbstractProgress of AI in biology is now being limited by the availability of high-quality biological sequence data from nature. To get past this data wall, we introduce BaseData, a new biological database built on top of a global, scalable biodiscovery pipeline. As of 2024, BaseData had already expanded the known protein universe by over 10x.

genomics↗