Search bioRxiv⌕ Search

Biology subjects

Rodriguez-Mier, P.

Publications and source records attributed to Rodriguez-Mier, P..

2 recordsLinked to original sources

Pathway analysis in metabolomics: pitfalls and best practice for the use of over-representation analysis

Over-representation analysis (ORA) is one of the commonest pathway analysis approaches used for the functional interpretation of metabolomics datasets. Despite the widespread use of ORA in metabolomics, the community lacks guidelines detailing its best-practice use. Many factors have a pronounced impact on the results, but to date their effects have received little systematic attention in the field. We developed in-silico simulations using five publicly available datasets and illustrated that changes in parameters, such as the background set, differential metabolite selection methods, and pathway database choice, could all lead to profoundly different ORA results. The use of a non-assay-specific background set, for example, resulted in large numbers of false-positive pathways. Pathway database choice, evaluated using three of the most popular metabolic pathway databases: KEGG, Reactome, and BioCyc, led to vastly different results in both the number and function of significantly enriched pathways. Metabolomics data specific factors, such as reliability of compound identification and assay chemical bias also impacted ORA results. Simulated metabolite misidentification rates as low as 4% resulted in both gain of false-positive pathways and loss of truly significant pathways across all datasets. Our results have several practical implications for ORA users, as well as those using alternative pathway analysis methods. We offer a set of recommendations for the use of ORA in metabolomics, alongside a set of minimal reporting guidelines, as a first step towards the standardisation of pathway analysis in metabolomics. Author summaryMetabolomics is a rapidly growing field of study involving the profiling of small molecules within an organism. It allows researchers to understand the effects of biological status (such as health or disease) on cellular biochemistry, and has wide-ranging applications, from biomarker discovery and personalised medicine in healthcare to crop protection and food security in agriculture. Pathway analysis helps to understand which biological pathways, representing collections of molecules performing a particular function, are involved in response to a disease phenotype, or drug treatment, for example. Over-representation analysis (ORA) is perhaps the most common pathway analysis method used in the metabolomics community. However, ORA can give drastically different results depending on the input data and parameters used. In this work, we have established the effects of these factors on ORA results using computational simulations applied to five real-world datasets. Based on our results, we offer the research community a set of best-practice recommendations applicable not only to ORA but also to other pathway analysis methods to help ensure the reliability and reproducibility of results.

bioinformatics↗

Diversity-based enumeration of optimal context-specific metabolic networks

The correct identification of metabolic activity in tissues or cells under different environmental or genetic conditions can be extremely elusive due to mechanisms such as post-transcriptional modification of enzymes or different rates in protein degradation, making difficult to perform predictions on the basis of gene expression alone. Context-specific metabolic network reconstruction can overcome these limitations by leveraging the integration of multi-omics data into genome-scale metabolic networks (GSMN). Using the experimental information, context-specific models are reconstructed by extracting from the GSMN the sub-network most consistent with the data, subject to biochemical constraints. One advantage is that these context-specific models have more predictive power since they are tailored to the specific organism and condition, containing only the reactions predicted to be active in such context. A major limitation of this approach is that the available information does not generally allow for an unambiguous characterization of the corresponding optimal metabolic sub-network, i.e., there are usually many different sub-network that optimally fit the experimental data. This set of optimal networks represent alternative explanations of the possible metabolic state. Ignoring the set of possible solutions reduces the ability to obtain relevant information about the metabolism and may bias the interpretation of the true metabolic state. In this work, we formalize the problem of enumeration of optimal metabolic networks, we implement a set of techniques that can be used to enumerate optimal networks, and we introduce DEXOM, a novel strategy for diversity-based extraction of optimal metabolic networks. Instead of enumerating the whole space of optimal metabolic networks, which can be computationally intractable, DEXOM samples solutions from the set of optimal metabolic sub-networks maximizing diversity in order to obtain a good representation of the possible metabolic state. We evaluate the solution diversity of the different techniques using simulated and real datasets, and we show how this method can be used to improve in-silico gene essentiality predictions in Saccharomyces Cerevisiae using diversity-based metabolic network ensembles. Both the code and the data used for this research are publicly available on GitHub1.

bioinformatics↗