Search bioRxiv⌕ Search

Biology subjects

Sidorenko, D.

Publications and source records attributed to Sidorenko, D..

4 recordsLinked to original sources

The End of Aging Clocks: Training Foundation Models to Reason in Aging and Longevity

The aging clock paradigm has yielded dozens of specialist models that can estimate chronological age or mortality from virtually any biodata type. Yet each such model operates within a fixed modality, relies on a predetermined feature set, and produces limited biological interpretation. Here, we report Longevity-LLM v0.1, a Qwen3-14B model fine-tuned through supervised and reinforcement learning regimes on DNA methylation, proteomics, clinical biomarker, and RNA expression data. Longevity-LLM achieves high ranks in the recently announced Longevity Bench, including such tasks as cancer survival and RNA- or proteome-based age prediction. After reinforcement fine-tuning, the model achieved a 4.34-year MAE in epigenetic age prediction, surpassing the Horvath multi-tissue clock. In addition to age prediction, Longevity-LLM can carry out numerous other tasks, including proteomic profile generation, for which it significantly outperforms all frontier LLMs. These results demonstrate that a single modestly sized LLM can match or replace purpose-built aging clocks across data modalities. This work constitutes an interim report from the initial sprint of our Multi-Modal AI Gym for Science (MMAI), an initiative dedicated to building foundation models for drug discovery and aging research.

bioinformatics↗

Longevity Bench: Are SotA LLMs ready for aging research?

Aging is a core biological process observed in most species and tissues, which is studied with a vast array of technologies. We argue that the abilities of AI systems to emulate aging and to accurately interpret biodata in its context are the key criteria to judge an LLMs utility in biomedical research. Here, we present LongevityBench -- a collection of tasks designed to assess whether foundation models grasp the fundamental principles of aging biology and can use low-level biodata to arrive at phenotype-level conclusions. The benchmark covers a variety of prediction targets including human time-to-death, mutations effect on lifespan, and age-dependent omics patterns. It spans all common biodata types used in longevity research: transcriptomes, DNA methylation profiles, proteomes, genomes, clinical blood tests and biometrics, as well as natural language annotations. After ranking state-of-the-art foundation models using LongevityBench, we highlight their weaknesses and outline procedures to maximize their utility in aging research and life sciences.

bioinformatics↗

DORA AI Scientist: Multi-agent Virtual Research Team for Scientific Exploration Discovery and Automated Report Generation

Modern goal-oriented scientific research process involves hierarchical teams of researchers of diverse backgrounds performing generalist and domain-specific tasks. Many of these tasks include hypothesis generation, literature review, data collection, cleanup, processing and analysis, experimental design, virtual and physical experiments, research report and academic paper writing, reference management, bibliography and quality control. Most of these tasks can be performed automatically or in a co-pilot mode by the generative reinforcement learning systems. In this paper, we introduce a versatile multi-agent scientific exploration and draft outline research assistant (DORA), which provides multiple templates and workflows for automated or semi-automated research studies and report generation. Under user guidance, it employs hierarchical teams of AI agents based on the plug-and-play generalist and domain-specific large language models (LLMs) exploiting a variety of specialized research tools and open data repositories and generates high-quality research outputs publication drafts with maximally-accurate references. DORA is designed to minimize the time and effort required for manuscript preparation, thereby enabling researchers to devote more attention to high-value discovery tasks. The system is constantly evolving with user feedback with regular feature and resource updates. The platform is available at https://dora.insilico.com.

bioinformatics↗

Precious3GPT: Multimodal Multi-Species Multi-Omics Multi-Tissue Transformer for Aging Research and Drug Discovery

We present a multimodal multi-species multi-omics multi-tissue transformer for aging research and drug discovery capable of performing multiple tasks such as age prediction across species, target discovery, tissue, sex, and disease sample classification, drug sensitivity prediction, replication of omics response and prediction of biological and phenotypic response to compound treatment. This model combines textual, tabular, and knowledge graph-derived representations of biological experiments to provide insights into molecular-level biological processes. We demonstrate that P3GPT has developed an intuition for the interactions between compounds, pathologies, and gene regulation in the context of multiple species and tissues. In these areas, it outperforms existing LLMs and we highlight its utility in diverse case studies. P3GPT is a general model that may be used as a target identification tool, aging clock, digital laboratory, and scientific assistant. The model is intended as a community resource available open source as well as via a Discord server.

bioinformatics↗