Search bioRxiv⌕ Search

Biology subjects

Sajeevan, K. A.

Publications and source records attributed to Sajeevan, K. A..

5 recordsLinked to original sources

Robust Prediction of Enzyme Variant Kinetics with RealKcat

Predicting enzyme kinetics directly from sequence remains a central challenge in computational biology, particularly in resolving the effects of mutations at catalytically essential residues. Existing models frequently overlook the functional consequences of such perturbations, often defaulting to wild-type predictions even in cases of substantial activity loss, thereby limiting their reliability for enzyme design and mechanistic inference. Here, we introduce RealKcat, a machine learning framework trained on KinHub-27k, a rigorously curated dataset of 27,176 experimentally reported enzyme-substrate entries consolidated from BRENDA, SABIO-RK, and UniProt and verified across 2,158 primary sources. To ensure biochemical realism, kinetic parameters were collapsed into order-of-magnitude bins, enabling predictions that are tolerant to experimental noise yet sensitive to functional shifts. RealKcat integrates ESM embeddings for enzyme sequences with ChemBERTa embeddings of affiliated substrate, producing a unified feature space of the chemical conversion that supports robust multi-class classification of both catalytic turnover (kCat) and substrate affinity (KM). Across cross-validation, hold-out, out-of-distribution, and few-shot evaluations--including a dense mutational landscape of alkaline phosphatase (PafA)--RealKcat consistently capturead the direction and magnitude of mutation-induced changes, while preserving discrimination in both wild-type and mutant contexts. Importantly, structural descriptors were deliberately excluded, as naive integration of structural features has been shown to impair model generalization, underscoring the primacy of rigorous dataset curation, biologically informed task formulation, and balanced evaluation metrics. RealKcat establishes a scalable and mutation-sensitive framework for enzyme kinetics prediction, offering a biologically grounded platform for enzyme engineering, metabolic modeling, and therapeutic design. Significance StatementEnzymes catalyze biochemical reactions that sustain life, and accurate measurement of their efficiency--expressed through turnover number (kCat) and substrate affinity (KM)--is fundamental to biotechnology, synthetic biology, and even pharmaceutical innovation. Yet experimental assays remain prohibitive, time-intensive, and sensitive to conditions such as pH, temperature, and ionic strength of the assay buffer, while existing computational approaches often lack sensitivity to catalytic-site mutations and are constrained by inconsistencies in public databases. RealKcat addresses these gaps by introducing a rigorously curated dataset (KinHub-27k) derived from manual review of 2,158 articles and augmented with 5,278 synthetic catalytic variants generated through alanine substitution at annotated catalytic residues. Leveraging protein and substrate embeddings and a classification scheme based on order-of-magnitude kinetic bins, RealKcat achieves state-of-the-art functional e-accuracy and, critically, demonstrates sensitivity to catalytic perturbations. By adopting e-accuracy--a performance metric that evaluates predictions within {+/-}1 order of magnitude, aligning with the practical utility of enzyme kinetics--RealKcat provides biologically meaningful assessments that conventional metrics often obscure. This work establishes a robust, mutation-aware predictive platform that advances computational enzyme design and extends applicability to biomanufacturing, metabolic engineering, and precision medicine.

systems biology↗

Improved Functional Classification of Hydrolases through Pairwise Structural Similarity of Reaction Cores

We report a systematic pipeline is for extracting the catalytically relevant reactive site in addition to the surrounding allosterically linked residue shells around the reaction site of the most diverse enzyme class - hydrolases, with known experimental structures. We first successfully extract 40196 such hydrolase reaction cores (RC) and collates them into a publicly accessible reaction core collection (RC-Hydrolase). We perform 128M pairwise shape comparison across RC-Hydrolase using a three-dimensional search engine and present 155,329 pair instances clustering them by 60% or higher similarities in a publicly available, visually interactive dataset. Robustness of defined RCs is shown to successfully capture experimentally known function-enhancing mutations distal to the active site in PETases. Allowing comparisons of enzyme reaction centers across functional spaces (ligands bound, EC classification numbers, and expression hosts) enables identification of enzyme backbones which can be minimally mutated to accommodate more than one type of catalytic activity thereby aiding rational design of multifunctional enzymes. We also demonstrate how such versatile enzyme backbones could be leveraged by the latest diffusion-based protein design models to design bespoke libraries of small molecule inhibitors, and structurally stable multifunctional enzyme pockets. With only sporadic successes in multifunctional enzyme design thus far, we provide strong structural priors for machine-learning-guided advanced enzyme engineering in the future.

bioinformatics↗

Exploring putative enteric methanogenesis inhibitors using molecular simulations and a graph neural network

Atmospheric methane (CH4) acts as a key contributor to global warming. As CH4 is a short-lived climate forcer (12 years atmospheric lifespan), its mitigation represents the most promising means to address climate change in the short term. Enteric CH4 (the biosynthesized CH4 from the rumen of ruminants) represents 5.1% of total global greenhouse gas (GHG) emissions, 23% of emissions from agriculture, and 27.2% of global CH4 emissions. Therefore, it is imperative to investigate methanogenesis inhibitors and their underlying modes of action. We hereby elucidate the detailed biophysical and thermodynamic interplay between anti-methanogenic molecules and cofactor F430 of methyl coenzyme M reductase and interpret the stoichiometric ratios and binding affinities of sixteen inhibitor molecules. We leverage this as prior in a graph neural network to first functionally cluster these sixteen known inhibitors among [~]54,000 bovine metabolites. We subsequently demonstrate a protocol to identify precursors to and putative inhibitors for methanogenesis, based on Tanimoto chemical similarity and membrane permeability predictions. This work lays the foundation for computational and de novo design of inhibitor molecules that retain/ reject one or more biochemical properties of known inhibitors discussed in this study. COVER ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=146 SRC="FIGDIR/small/613350v1_ufig1.gif" ALT="Figure 1"> View larger version (49K): org.highwire.dtl.DTLVardef@1fd4518org.highwire.dtl.DTLVardef@c34d87org.highwire.dtl.DTLVardef@16eef9org.highwire.dtl.DTLVardef@1a33512_HPS_FORMAT_FIGEXP M_FIG C_FIG

molecular biology↗

Dissecting Metabolic Landscape of Alveolar Macrophage

The highly plastic nature of Alveolar Macrophage (AM) plays a crucial role in the defense against inhaled particulates and pathogens in the lungs. Depending upon the signal, AM acquires either classically activated M1 phenotype or alternatively activated M2 phenotype. These phenotypes have specific functions and unique metabolic traits such as upregulated glycolysis and pentose phosphate pathway in M1 phase and enhanced oxidative phosphorylation and tricarboxylic acid cycle during M2 phase that help maintain the sterility of the lungs. In this study, we investigate the metabolic shift in the activated phases of AM (M1 and M2 phase) and highlight the roles of pathways other than the typical players of central carbon metabolism. Pathogenesis is a complex and elongated process where the heightened requirement for energy is matched by metabolic shifts that supplement immune response and maintain homeostasis. The first step of pathogenesis is fever; however, analyzing the role of physical parameters such as temperature is challenging. Here, we observe the effect of an increase in temperature on pathways such as glycolysis, pentose phosphate pathway, oxidative phosphorylation, tricarboxylic acid cycle, amino acid metabolism, and leukotriene metabolism. We report the role of temperature as a catalyst to the immune response of the cell. The activity of pathways such as pyruvate metabolism, arachidonic acid metabolism, chondroitin/heparan sulfate biosynthesis, and heparan sulfate degradation are found to be important driving forces in the M1/M2 phenotype. We have also identified a list of 34 reactions such as nitric oxide production from arginine and the conversion of glycogenin to UDP which play major roles in the metabolic models and prompt the shift of the M2 phenotype to M1 and vice versa. In future, these reactions could further be probed as major contributors in designing effective therapeutic targets against severe respiratory diseases. Author SummaryAlveolar macrophage (AM) is highly plastic in nature and has a wide range of functions including invasion/killing of bacteria to maintaining the homeostasis in the lungs. The regulatory mechanism involved in the alveolar macrophage polarization is essential to fight against severe respiratory conditions (pathogens and particulates). Over the years, experiments on mouse/rat models have been used to draw insightful inferences. However, recent advances have highlighted the lack of transmission from non-human models to successful in vivo human experiments. Hence using genome-scale metabolic (GSM) models to understand the unique metabolic traits of human alveolar macrophages and comprehend the complex metabolic underpinnings that govern the polarization can lead to novel therapeutic strategies. The GSM models of AMs thus far, has not incorporated the activated phases of AM. Here, we aim to exhaustively dissect the metabolic landscape and capabilities of AM in its healthy and activated stages. We carefully explore the changes in reaction fluxes under each of the conditions to understand the role and function of all the pathways with special attention to pathways away from central carbon metabolism. Understanding the characteristics of each phase of AM has applications that could help improve the therapeutic approaches against respiratory conditions.

systems biology↗

Multi-organ Metabolic Model of Zea mays Connects Temperature Stress with Thermodynamics-Reducing Power-Energy Generation Axis

Global climate change has severely impacted maize productivity. A holistic understanding of metabolic crosstalk among its organs is essential to address this issue. Thus, we reconstructed the first multi-organ maize genome-scale metabolic model, iZMA6517, and contextualized it with heat and cold stress-related transcriptomics data using the novel EXpression disTributed REAction flux Measurement (EXTREAM) algorithm. Furthermore, implementing metabolic bottleneck analysis on contextualized models revealed fundamental differences between these stresses. While both stresses had reducing power bottlenecks, heat stress had additional energy generation bottlenecks. To tie these signatures, we performed thermodynamic driving force analysis, revealing thermodynamics-reducing power-energy generation axis dictating the nature of temperature stress responses. Thus, for global food security, a temperature-tolerant maize ideotype can be engineered by leveraging the proposed thermodynamics-reducing power-energy generation axis. We experimentally inoculated maize root with a beneficial mycorrhizal fungus, Rhizophagus irregularis, and as a proof of concept demonstrated its potential to alleviate temperature stress. In summary, this study will guide the engineering effort of temperature stress-tolerant maize ideotypes.

plant biology↗