Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.06.03.730004

Comprehensive evaluation of LLM capabilities for interpretation and analysis of genome-scale metabolic models in metabolic engineering

Abstract

Genome-scale metabolic models (GSMs) underpin pathway and strain engineering by linking genes to metabolic reactions and enabling system-level simulation of cellular fluxes and intervention effects, yet end-to-end analysis workflows remain fragmented, expert-demanding, and slow to adapt. Large language models (LLMs) could transform this landscape, lowering the barrier by explaining concepts, interpreting GSM files, and turning natural-language instructions into valid analysis code, thereby substantially mitigating the time, effort, and expertise required. However, their reliability for domain-specific tasks remains unexplored. Here, we delivered a systematic benchmark of four leading LLMs (GPT-4, Gemini, Claude, DeepSeek-R1) across four task areas central to metabolic engineering: domain knowledge, metabolic flux prediction, pathway construction, and flux optimization. For benchmarking, we introduced a standardized, rubric-based evaluation framework that uses multi-LLM automated scoring (an ensemble of LLM-as-a-judge assessments) and two distinct sets of nine task-tailored metrics (domain vs coding-focused tasks), rated on a 1-5 scale (up to 45 per task), covering scientific validity and code executability where applicable. Across tasks, we reveal consistent strengths (conceptual explanation, code synthesis) and critical failure modes (e.g., context window limitations, incorrect identifier assumptions, strain-dependent reasoning errors, and errors in domain-specific algorithms). In aggregate, DeepSeek-R1 led in domain tasks, narrowly edging GPT-4, Claude, and Gemini, demonstrating that conceptual biological logic remains highly invariant across architectures. In contrast, Gemini achieved the highest score for coding tasks, distinguished by functional execution and excelled in error handling, documentation, and readability, followed by GPT-4, Claude, and DeepSeek. We also evaluated LLM self-inspection capability by injecting subtle, consequential faults: a stoichiometric sign error causing mass imbalance and an omitted pathway reaction. We reveal that conversational "blind search" prompting completely fails to localize these network faults. Instead, robust error localization requires prompts reframed with domain-informed constraints that force the LLM to leverage tool-assisted code procedures, such as COBRApy mass-balance functions. Together, this work establishes an evidence-based baseline for LLM-enabled GSM analysis, providing actionable guidance for building reliable, automation-ready workflows for pathway and strain design. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=88 SRC="FIGDIR/small/730004v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@413e29org.highwire.dtl.DTLVardef@1581de2org.highwire.dtl.DTLVardef@11ef7borg.highwire.dtl.DTLVardef@1817653_HPS_FORMAT_FIGEXP M_FIG C_FIG

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yeoh, J. W., Patro, C. P. K., Wong, L., Poh, C. L.. 2026-06-08. Comprehensive evaluation of LLM capabilities for interpretation and analysis of genome-scale metabolic models in metabolic engineering. https://doi.org/10.64898/2026.06.03.730004

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Limit-pushing overexpression reveals constraints on protein abundance

Proteins are often classified as toxic or non-toxic without measuring the abundance reached, leaving constraints on tolerable protein abundance unresolved. We established a limit-pushing approach in Saccharomyces cerevisiae combining strong inducible expression with gTOW-mediated high-copy selection to counteract copy-number compensation while measuring protein abundance and growth. Nearly all of approximately 80 chromosome I proteins severely inhibited growth or reduced viability at sufficiently high abundance. We established IE50, the expression level associated with a 50% reduction in growth rate, to quantify their widely varying overexpression tolerance. IE50 was positively associated with predicted structural order and cytoplasmic localization propensity and negatively associated with sulphur content. Single-cell imaging linked higher tolerance to proteins remaining cytoplasmic without becoming aggregation-positive and revealed abundance-dependent changes in localization and organelle morphology. At extreme abundance, Fun12, Nup60, and Pex22 generated distinct large-scale intracellular states through specific sequence regions. These findings establish overexpression toxicity as a quantitative property linked to protein characteristics and reveal both constraints on tolerable abundance and sequence-dependent capacities for intracellular organization.

systems biology↗

Accessing Enzyme Kinetic Data and Prediction Methods at Scale

Enzyme kinetic parameters inform metabolic models, yet experimental measurements are sparse. A growing body of work predicts them from protein and substrate features, but software fragmentation hinders adoption, so downstream tools lock into the most accessible method. We present OpenKinetics Predictor (at predictor.openkinetics.org), an open-source platform integrating thirteen methods in isolated environments behind one interface. The platform optionally reports similarity between query proteins and each method's training data to contextualise reliability. A common featurisation-prediction abstraction keeps it extensible, and independent parties, including original authors, contributed many methods. We pair it with a data portal (at data.openkinetics.org) that exposes CatLog, a curated kinetic dataset, with precomputed embeddings, predicted binding sites, and standardised splits. Both offer a web interface and an API, and the GECKO modelling toolbox already calls the predictor API. As a case study, we predict across an E. coli model and find inter-predictor agreement varies with metabolic context and data availability.

systems biology↗

A thermoregulatory design principle for transitions into hypometabolism

Mammals entering torpor or hibernation undergo an abrupt transition from normothermia to hypothermia, yet how thermoregulation enables this switch remains poorly understood. Here, we identify dynamical signatures that precede these transitions and a mathematical principle that can generate them. In fasting-induced torpor in mice, body-temperature fluctuations increased before torpor onset, providing an early-warning signal that tracked proximity to the transition better than temperature decline alone. A heat-balance model showed that reducing how strongly the effective heat-loss coefficient depends on body temperature reorganizes thermoregulatory stability, allowing a low-temperature equilibrium to emerge while the normothermic state remains stable. This organization is consistent with a symmetry-broken pitchfork involving a saddle-node. Similar increases in temperature fluctuations preceded hibernation onset in hamsters. These findings link pre-transition temperature dynamics to changes in the underlying thermoregulatory landscape and provide a framework for detecting and understanding transitions from normothermia to hypothermia.

systems biology↗