Search bioRxiv⌕ Search

Biology subjects

Elizarraras, J. M.

Publications and source records attributed to Elizarraras, J. M..

2 recordsLinked to original sources

XL-Ranker: A Computational Workflow for Prioritizing Protein-Protein Interactions from Cross-Linking Mass Spectrometry Data

Protein-protein interactions (PPIs) are central to virtually all biological processes, and their disruption can lead to a wide spectrum of human diseases. Cross-linking mass spectrometry (XL-MS) enables proteome-scale detection of interacting peptide pairs, from which PPIs can be inferred. However, interpretating XL-MS data to assign homologous cross-linked peptide pairs to specific protein interactions can be challenging. A major hurdle arises when a single peptide can map to multiple proteins, and when two such peptides are paired, the number of possible protein-protein interactions increases combinatorially. This mapping ambiguity can lead to inflated interaction networks that could compromise downstream analysis and biological interpretation. However, this problem has often been overlooked or addressed using heuristic approaches in previous XL-MS studies. To tackle this challenge, we developed XL-Ranker, a computational framework that combines a set cover graph algorithm and machine learning to systematically resolve peptide-mapping ambiguity and infer high-confidence PPIs from XL-MS data. Applied to an XL-MS dataset from HEK293 cells, XL-Ranker identified a high-confidence network with 880 PPIs involving 964 unique genes. Among all possible PPIs, the ones selected by XL-Ranker for inclusion in the final network had significantly higher interaction scores in the STRING database than the excluded ones. Network analysis further demonstrated that these interactions form biologically meaningful clusters, supporting the accuracy of our approach. In summary, XL-Ranker provides a practical solution to a key analytical challenge in XL-MS data interpretation, enhancing the reliability of PPI discovery.

bioinformatics↗

MetaSage: Machine Learning-Based Prioritization of Metabolic Regulators from Multi-Omics Data

Dysregulation of metabolites is a hallmark of cancer, yet the underlying regulatory mechanisms remain poorly understood. To systematically explore metabolic regulation across cancers, we developed an XGBoost-based machine learning pipeline, MetaSage, that integrates context-agnostic knowledge graph with multi-omics datasets. Using harmonized data from 15 cohorts spanning 11 cancer types, we identified 442 variable metabolites and found that both genes and upstream metabolites showed comparable regulatory influence. Predictable metabolites, defined by a significant correlation between predicted and measured levels, were identified using our pipeline and varied widely across cohorts-partially due to the batch effect. For each predictable metabolite, key regulatory features were determined using Shapley values. This yielded 1,146 gene features and 363 precursor metabolites as important regulators. Network analysis of 22 recurrent metabolites revealed a mix of conserved and cancer type-specific regulatory patterns. Our framework enables robust discovery of metabolite regulation and therapeutic insights in cancer.

bioinformatics↗