Search bioRxivSearch

bioRxiv · 10.1101/077636

Managing Uncertainty in Metabolic Network Structure and Improving Predictions Using EnsembleFBA

Abstract

Genome-scale metabolic network reconstructions (GENREs) are repositories of knowledge about the metabolic processes that occur in an organism. GENREs have been used to discover and interpret metabolic functions, and to engineer novel network structures. A major barrier preventing more widespread use of GENREs, particularly to study non-model organisms, is the extensive time required to produce a high-quality GENRE. Many automated approaches have been developed which reduce this time requirement, but automatically-reconstructed draft GENREs still require curation before useful predictions can be made. We present a novel ensemble approach to the analysis of GENREs which improves the predictive capabilities of draft GENREs and is compatible with many automated reconstruction approaches. We refer to this new approach as Ensemble Flux Balance Analysis (EnsembleFBA). We validate EnsembleFBA by predicting growth and gene essentiality in the model organism Pseudomonas aeruginosa UCBPP-PA14. We demonstrate how EnsembleFBA can be included in a systems biology workflow by predicting essential genes in six Streptococcus species and mapping the essential genes to small molecule ligands from DrugBank. We found that some metabolic subsystems contribute disproportionately to the set of predicted essential reactions in a way that is unique to each Streptococcus species. These species-specific network structures lead to species-specific outcomes from small molecule interactions. Through these analyses of P. aeruginosa and six Streptococci, we show that ensembles increase the quality of predictions without drastically increasing reconstruction time, thus making GENRE approaches more practical for applications which require predictions for many non-model organisms. All of our functions and accompanying example code are available in an open online repository.\n\nAuthor SummaryMetabolism is the driving force behind all biological activity. Genome-scale metabolic network reconstructions (GENREs) are representations of metabolic systems that can be analyzed mathematically to make predictions about how a biochemical system will behave as well as to design biochemical systems with new properties. GENREs have traditionally been reconstructed manually, which can require extensive time and effort. Recent software solutions automate the process (drastically reducing the required effort) but the resulting GENREs are of lower quality and produce less reliable predictions than the manually-curated versions. We present a novel method (\"EnsembleFBA\") which overcomes uncertainties involved in automated reconstruction by pooling many different draft GENREs together into an ensemble. We tested EnsembleFBA by predicting the growth and essential genes of the common pathogen Pseudomonas aeruginosa. We found that when predicting growth or essential genes, ensembles of GENREs achieved much better precision or captured many more essential genes than any of the individual GENREs within the ensemble. By improving the predictions that can be made with automatically-generated GENREs, we open the door to studying systems which would otherwise be infeasible.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Matthew Biggs, Jason A Papin. 2016-09-26. Managing Uncertainty in Metabolic Network Structure and Improving Predictions Using EnsembleFBA. https://doi.org/10.1101/077636

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Single Cell Phenotyping Reveals Heterogeneity among Haematopoietic Stem Cells Following Infection

The haematopoietic stem cell (HSC) niche provides essential micro-environmental cues for the production and maintenance of HSCs within the bone marrow. During inflammation, haematopoietic dynamics are perturbed, but it is not known whether changes to the HSC-niche interaction occur as a result. We visualise HSCs directly in vivo, enabling detailed analysis of the 3D niche dynamics and migration patterns in murine bone marrow following Trichinella spiralis infection. Spatial statistical analysis of these HSC trajectories reveals two distinct modes of HSC behaviour: (i) a pattern of revisiting previously explored space, and (ii) a pattern of exploring new space. Whereas HSCs from control donors predominantly follow pattern (i), those from infected mice adopt both strategies. Using detailed computational analyses of cell migration tracks and life-history theory, we show that the increased motility of HSCs following infection can, perhaps counterintuitively, enable mice to cope better in deteriorating HSC-niche micro-environments following infection.\n\nAuthor SummaryHaematopoietic stem cells reside in the bone marrow where they are crucially maintained by an incompletely-determined set of niche factors. Recently it has been shown that chronic infection profoundly affects haematopoiesis by exhausting stem cell function, but these changes have not yet been resolved at the single cell level. Here we show that the stem cell-niche interactions triggered by infection are heterogeneous whereby cells exhibit different behavioural patterns: for some, movement is highly restricted, while others explore much larger regions of space over time. Overall, cells from infected mice display higher levels of persistence. This can be thought of as a search strategy: during infection the signals passed between stem cells and the niche may be blocked or inhibited. Resultantly, stem cells must choose to either cling on, or to leave in search of a better environment. The heterogeneity that these cells display has immediate consequences for translational therapies involving bone marrow transplant, and the effects that infection might have on these procedures.

Systems Biology

Analysis of noise mechanisms in cell size control

At the single-cell level, noise features in multiple ways through the inherent stochasticity of biomolecular processes, random partitioning of resources at division, and fluctuations in cellular growth rates. How these diverse noise mechanisms combine to drive variations in cell size within an isoclonal population is not well understood. To address this problem, we systematically investigate the contributions of different noise sources in well-known paradigms of cell-size control, such as the adder (division occurs after adding a fixed size from birth) and the sizer (division occurs upon reaching a size threshold). Analysis reveals that variance in cell size is most sensitive to errors in partitioning of volume among daughter cells, and not surprisingly, this process is well regulated among microbes. Moreover, depending on the dominant noise mechanism, different size control strategies (or a combination of them) provide efficient buffering of intercellular size variations. We further explore mixer models of size control, where a timer phase precedes/follows an adder, as has been proposed in Caulobacter crescentus. While mixing a timer with an adder can sometimes attenuate size variations, it invariably leads to higher-order moments growing unboundedly over time. This results in the cell size following a power-law distribution with an exponent that is inversely dependent on the noise in the timer phase. Consistent with theory, we find evidence of power-law statistics in the tail of C. crescentus cell-size distribution, but there is a huge discrepancy in the power-law exponent as estimated from data and theory. However, the discrepancy is removed after data reveals that the size added by individual newborns from birth to division itself exhibits power-law statistics. Taken together, this study provides key insights into the role of noise mechanisms in size homeostasis, and suggests an inextricable link between timer-based models of size control and heavy-tailed cell size distributions.

Systems Biology

A mathematical approach for secondary structure analysis can provide an eyehole to the RNA world

The RNA pseudoknot is a conserved secondary structure encountered in a number of ribozymes, which assume a central role in the RNA world hypothesis. However, RNA folding algorithms could not predict pseudoknots until recently. Analytic combinatorics - a newly arisen mathematical field - has introduced a way of enumerating different RNA configurations and quantifying RNA pseudoknot structure robustness and evolvability, two features that drive their molecular evolution. I will present a mathematicians viewpoint of RNA secondary structures, and explain how analytic combinatorics applied on RNA sequence to structure maps can represent a valuable tool for understanding RNA secondary structure evolution. Analytic combinatorics can be implemented for the optimization of RNA secondary structure prediction algorithms, the derivation of molecular evolution mathematical models, as well as in a number of biotechnological applications, such as biosensors, riboswitches etc. Moreover, it showcases how the integration of biology and mathematics can provide a different viewpoint into the RNA world.

Systems Biology