Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 847 records · Page 47Linked to original sources

COBRAme: A Computational Framework for Building and Manipulating Models of Metabolism and Gene Expression

Genome-scale models of metabolism and macromolecular expression (ME-models) explicitly compute the optimal proteome composition of a growing cell. ME-models expand upon the well-established genome-scale models of metabolism (M-models), and they enable new and exciting insights that are fundamental to understanding the basis of cellular growth. ME-models have increased predictive capabilities and accuracy due to their inclusion of the biosynthetic costs for the machinery of life, but they come with a significant increase in model size and complexity. This challenge results in models which are both difficult to compute and challenging to understand conceptually. As a result, ME-models exist for only two organisms (Escherichia coli and Thermotoga maritima) and are still used by relatively few researchers. To address these challenges, we have developed a new software framework called COBRAme for building and simulating ME-models. It is coded in Python and built on COBRApy, a popular platform for using M-models. COBRAme streamlines computation and analysis of ME-models. It provides tools to simplify constructing and editing ME-models to enable ME-model reconstructions for new organisms. We used COBRAme to reconstruct a condensed E. coli ME-model called iJL1678b-ME. This reformulated model gives virtually identical solutions to previous E. coli ME-models while using [1/4] the number of free variables and solving in less than 10 minutes, a marked improvement over the 6 hour solve time of previous ME-model formulations. This manuscript outlines the architecture of COBRAme and demonstrates how ME-models can be reconstructed and edited most efficiently using the software.

systems biology

Reduced protein expression in a virus attenuated by codon deoptimization

The engineering of hundreds of synonymous codon changes into a viral genome appears to provide a general means of achieving attenuation. The mechanistic underpinnings of this approach remain enignmatic, however. Using quantitative proteomics and RNA sequencing, we explore the molecular basis of attenuation in a strain of bacteriophage T7 whose major capsid gene was engineered to carry 182 suboptimal codons. As expected, there was no evident effect of the recoding on transcription. Proteomic observations revealed that translation is halved for the recoded major capsid gene, and a smaller reduction applies to a few genes downstream, potentially caused by translational coupling. Viral burst size is also approximately halved, and the fitness drop accompanying attenuation is compatible with the reduced burst size. Overall, the fitness effect and molecular basis of attenuation by codon deoptimization are compatible with a relatively simple model of reduced translation of a few genes and a consequent diminished virion assembly. This mechanism is simpler than that operating in eukaryotic viruses.

systems biology

Topological and Kinetic Determinants of the Modal Matrices of Dynamic Models of Metabolism

Linear analysis of kinetic models of metabolism can help in understanding the dynamic response of metabolic networks. Central to linear analysis of these models are two key matrices: the Jacobian matrix (J) and its modal matrix (M-1). The modal matrix contains dynamically independent motions of the kinetic model, and it is sparse in practice. Understanding the sparsity structure of the modal matrix provides insight into metabolic network dynamics. In this study, we analyze the relationship between J and M-1. First, we show that diagonal dominance occurs in a substantial fraction of the rows of J, resulting in simple modal structures within M-1. Dominant diagonal elements in J approximate the eigenvalues corresponding to these simple modes, in which a single metabolite is driven back to its reference state on a characteristic timescale. Second, we analyze more complicated mode structures in M-1, in which two or more variables move in a certain ratio relative to one another on defined time scales. We show that complicated modes originate from sub-matrices of topologically connected elements of similar magnitude in J. Third, we describe the origin of these mode structure features based on the network stoichiometric matrix S and the reaction kinetic gradient matrix G. We demonstrate that the topologically-connected reaction sensitivities of similar magnitude in G play a central role in determining the mode structure. Ratios of these reaction sensitivities represent equilibrium balances of half reactions that are defined by linearization of the bilinear mass action rate laws followed by enzymatic reactions. These half-reaction equilibrium ratios are key determinants of modal structure for both simple and complicated modes. The work presented here helps to establish a foundation for understanding the dynamics of kinetic models of metabolism, which are rooted in the network structure and the kinetic properties of reactions.

systems biology

A generator of morphological clones for plant species

Detailed and realistic tree form generators have numerous applications in ecology and forestry. Here, we present an algorithm for generating morphological tree \"clones\" based on the detailed reconstruction of the laser scanning data, statistical measure of similarity, and a plant growth algorithm with simple stochastic rules. The algorithm is designed to produce tree forms, i.e. morphological clones, similar as a whole (coarse-grain scale), but varying in minute details of organization (fine-grain scale). We present a general procedure for obtaining these morphological clones. Although we opted for certain choices in our algorithm, its various parts may vary depending on the application. Namely, we have shown that specific multi-purpose procedural stochastic growth model can be algorithmically adjusted to produce the morphological clones replicated from the target experimentally measured tree. For this, we have developed a statistical measure of similarity (structural distance) between any given pair of trees, which allows for the comprehensive comparing of the tree morphologies in question by means of empirical distributions describing geometrical and topological features of a tree. Our algorithm can be used in variety of applications and contexts for exploration of the morphological potential of the growth models, arising in all sectors of plant science research.\n\nSummary StatementWe present an algorithmic framework, based on the Bayesian inference, for generating morphological tree clones using a combination of stochastic growth models and experimentally derived tree structures.

systems biology

Systematic analysis of cell phenotypes and cellular social networks in tissues using the multiplexed image cytometry analysis toolbox (miCAT)

Single-cell, spatially resolved omics analysis of tissues is poised to transform biomedical research and clinical practice. We have developed an open-source computational histology topography analysis toolbox (histoCAT) to enable the interactive, quantitative, and comprehensive exploration of phenotypes of individual cells, cell-to-cell interactions, microenvironments, and morphological structures within intact tissues. histoCAT will be useful in all areas of tissue-based research. We highlight the unique abilities of histoCAT by analysis of highly multiplexed mass cytometry images of human breast cancer tissues.\n\nTechnological advances in the multi-parametric analysis of single cells have revealed an unprecedented heterogeneity of cellular phenotypes and functional states that are concealed in population-based studies1-3. Each cellular phenotype is defined by the interplay of its internal state and the environment in which ...

systems biology

Bayesian inference of transcription dynamics from population snapshots of single-molecule RNA FISH in single cells

Single-molecule RNA fluorescence in situ hybridization (smFISH) provides unparalleled resolution on the abundance and localization of nascent and mature transcripts in single cells. Gene expression dynamics are typically inferred by measuring mRNA abundance in small numbers of fixed cells sampled from a population at multiple time-points after induction. The sparse data that arise from the small number of cells obtained using smFISH present a challenge for inferring transcription dynamics. Here, we developed a computational pipeline (BayFish) to infer kinetic parameters of gene expression from smFISH data at multiple time points after induction. Given an underlying model of gene expression, BayFish uses a Monte Carlo method to estimate the Bayesian posterior probability of the model parameters and quantify the parameter uncertainty given the observed smFISH data. We tested BayFish on smFISH measurements of the neuronal activity inducible gene Npas4 in primary neurons. We showed that a 2-state promoter model can recapitulate Npas4 dynamics after induction and we inferred that the transition rate from the promoter OFF state to the ON state is increased by the stimulus.\n\nAuthor SummaryGene expression can exhibit cell-to-cell variability due to the stochastic nature of biochemical reactions. Single cell assays (e.g. smFISH) directly quantify stochastic gene expression by measuring the number of active promoters and transcripts per cell in a population of cells. The data are distributions and their shape and time-evolution contain critical information on the underlying process of gene expression. Recent work has combined models of stochastic gene expression with maximum likelihood methods to infer kinetic parameters from smFISH distributions. However, these approaches do not provide a probability distribution or likelihood of model parameters inferred from the smFISH data. This information is useful because it indicates which parameters are loosely constrained by the data and suggests follow up experiments. We developed a suite of MATLAB programs (BayFish) that estimate the Bayesian posterior probability of model parameters from smFISH data. The user specifies an underlying model of stochastic gene expression with unknown parameters ({theta}) and provides smFISH data (Y). BayFish uses a Monte Carlo algorithm to estimate the Bayesian posterior probability P({theta}|Y) of model parameters. BayFish is easily modified and can be applied to other models of stochastic gene expression and smFISH data sets.

systems biology

Profiling metabolic flux modes by enzyme cost reveals variable trade-offs between growth and yield in Escherichia coli.

Microbes may maximize the number of daughter cells per time or per amount of nutrients consumed. These two strategies correspond, respectively, to the use of enzyme-efficient or substrate-efficient metabolic pathways. In reality, fast growth is often associated with wasteful, yield-inefficient metabolism, and a general thermodynamic trade-off between growth rate and biomass yield has been proposed to explain this. We studied growth rate/yield trade-offs by using a novel modeling framework, Enzyme-Flux Cost Minimization (EFCM) and by assuming that the growth rate depends directly on the enzyme investment per rate of biomass production. In a comprehensive mathematical model of core metabolism in E. coli, we screened all elementary flux modes leading to cell synthesis, characterized them by the growth rates and yields they provide, and studied the shape of the resulting rate/yield Pareto front. By varying the model parameters, we found that the rate/yield trade-off is not universal, but depends on metabolic kinetics and environmental conditions. A prominent trade-off emerges under oxygen-limited growth, where yield-inefficient pathways support a 2-to-3 times higher growth rate than yield-efficient pathways. EFCM can be widely used to predict optimal metabolic states and growth rates under varying nutrient levels, perturbations of enzyme parameters, and single or multiple gene knockouts.\n\nAuthor SummaryWhen cells compete for nutrients, those that grow faster and produce more offspring per time are favored by natural selection. In contrast, when cells need to maximize the cell number at a limited nutrient supply, fast growth does not matter and an efficient use of nutrients (i.e. high biomass yield) is essential. This raises a basic question about metabolism: can cells achieve high growth rates and yields simultaneously, or is there a conflict between the two goals? Using a new modeling method called Enzymatic Flux Cost Minimization (EFCM), we predict cellular growth rates and find that growth rate/yield trade-offs and the ensuing preference for enzyme-efficient or substrate-efficient metabolic pathways are not universal, but depend on growth conditions such as external glucose and oxygen concentrations.

systems biology

Constraint-based modeling identifies new putative targets to fight colistin-resistant A. baumannii infections.

Acinetobacter baumannii is a clinical threat to human health, causing major infection outbreaks worldwide. As new drugs against Gram-negative bacteria do not seem to be forthcoming, and due to the microbial capability of acquiring multi-resistance, there is an urgent need for novel therapeutic targets. Here we have derived a list of new potential targets by means of metabolic reconstruction and modelling of A. baumannii ATCC 19606. By integrating constraint-based modelling with gene expression data, we simulated microbial growth in normal and stressful conditions (i.e. following antibiotic exposure). This allowed us to describe the metabolic reprogramming that occurs in this bacterium when treated with colistin (the currently adopted last-line treatment) and identify a set of genes that are primary targets for developing new drugs against A. baumannii, including colistin-resistant strains. It can be anticipated that the metabolic model presented herein will represent a solid and reliable resource for the future treatment of A. baumannii infections.

systems biology

Active and poised promoter states drive folding of the extended HoxB locus in mouse embryonic stem cells

Gene expression states influence the three-dimensional conformation of the genome through poorly understood mechanisms. Here, we investigate the conformation of the murine HoxB locus, a gene-dense genomic region containing closely spaced genes with distinct activation states in mouse embryonic stem (ES) cells. To predict possible folding scenarios, we performed computer simulations of polymer models informed with different chromatin occupancy features, which define promoter activation states or CTCF binding sites. Single cell imaging of the locus folding was performed to test model predictions. While CTCF occupancy alone fails to predict the in vivo folding at genomic length scale of 10 kb, we found that homotypic interactions between active and Polycomb-repressed promoters co-occurring in the same DNA fibre fully explain the HoxB folding patterns imaged in single cells. We identify state-dependent promoter interactions as major drivers of chromatin folding in gene-dense regions.

systems biology

Revealing Age-Related Changes of Adult Hippocampal Neurogenesis

In the adult hippocampus, neural stem cells (NSCs) continuously produce new neurons that integrate into the neuronal network to modulate learning and memory. The amount and quality of newly generated neurons decline with age, which can be counteracted by increasing intrinsic Wnt activity in NSCs. However, the precise cellular changes underlying this age-related decline or its rescue through Wnt remain unclear. The present study combines development of a mathematical model and experimental data to address features controlling stem cell dynamics. We show that available experimental data fit a model in which quiescent NSCs can either become activated to divide or undergo depletion events, consisting of astrocytic transformation and apoptosis. Additionally, we demonstrate that aged NSCs remain longer in quiescence and have a higher probability to become re-activated versus being depleted. Finally, our model explains that high NSC-Wnt activity leads to longer time in quiescence while augmenting the probability of activation.

systems biology

Quantifying the entropic cost of cellular growth control

We quantify the amount of regulation required to control growth in living cells by a Maximum Entropy approach to the space of underlying metabolic states described by genome-scale models. Results obtained for E. coli and human cells are consistent with experiments and point to different regulatory strategies by which growth can be fostered or repressed. Moreover we explicitly connect the inverse temperature that controls MaxEnt distributions to the growth dynamics, showing that the initial size of a colony may be crucial in determining how an exponentially growing population organizes the phenotypic space.

systems biology

A yield-cost tradeoff governs Escherichia coli's decision between fermentation and respiration in carbon-limited growth

Many microbial systems are known to actively reshape their proteomes in response to changes in growth conditions induced e.g. by nutritional stress or antibiotics. Part of the re-allocation accounts for the fact that, as the growth rate is limited by targeting specific metabolic activities, cells simply respond by fine-tuning their proteome to invest more resources into the limiting activity (i.e. by synthesizing more proteins devoted to it). However, this is often accompanied by an overall re-organization of metabolism, aimed at improving the growth yield under limitation by re-wiring resource through different pathways. While both effects impact proteome composition, the latter underlies a more complex systemic response to stress. By focusing on E. coli's acetate switch, we use mathematical modeling and a re-analysis of empirical data to show that the transition from a predominantly fermentative to a predominantly respirative metabolism in carbon-limited growth results from the trade-off between maximizing the growth yield and minimizing its costs in terms of required the proteome share. In particular, E. coli's metabolic phenotypes appear to be Pareto-optimal for these objective functions over a broad range of dilutions.

systems biology

An indicator cell assay for blood-based diagnostics

We have established proof of principle for the Indicator Cell Assay Platform (iCAP), a broadly applicable tool for blood-based diagnostics that uses specifically-selected, standardized cells as biosensors, relying on their innate ability to integrate and respond to diverse signals present in patients blood. To develop an assay, indicator cells are exposed in vitro to serum from case or control subjects and their global differential response patterns are used to train reliable, cost-effective disease classifiers based on a small number of features. In a feasibility study, the iCAP detected pre-symptomatic disease in a murine model of amyotrophic lateral sclerosis (ALS) with 94% accuracy (p-Value=3.81E-6) and correctly identified samples from a murine Huntingtons disease model as non-carriers of ALS. In a preliminary human disease assay, the iCAP detected early stage Alzheimers disease with 72% cross-validated accuracy (p-Value=3.10E-3). For both assays, iCAP features were enriched for disease-related genes, supporting the assays relevance for disease research.

systems biology

MODELING SIGNALING-DEPENDENT PLURIPOTENT CELL STATES WITH BOOLEAN LOGIC CAN PREDICT CELL FATE TRANSITIONS

Pluripotent stem cells (PSCs) exist in multiple stable states, each with specific cellular properties and molecular signatures. The process by which pluripotency is either maintained or destabilized to initiate specific developmental programs is poorly understood. We have developed a model to predict stabilized PSC gene regulatory network (GRN) states in response to combinations of input signals. While previous attempts to model PSC fate have been limited to static cell compositions, our approach enables simulations of dynamic heterogeneity by combining an Asynchronous Boolean Simulation (ABS) strategy with simulated single cell fate transitions using Strongly Connected Components (SCCs). This computational framework was applied to a reverse-engineered and curated core GRN for mouse embryonic stem cells (mESCs) to simulate responses to LIF, Wnt/{beta}-catenin, FGF/ERK, BMP4, and Activin A/Nodal pathway activation. For these input signals, our simulations exhibit strong predictive power for gene expression patterns, cell population composition, and nodes controlling cell fate transitions. The model predictions extend into early PSC differentiation, demonstrating, for example, that a Cdx2-high/Oct4-low state can be efficiently and robustly generated from mESCs residing in a naive and signal-receptive state sustained by combinations of signaling activators and inhibitors.\n\nOne Sentence SummaryPredictive control of pluripotent stem cell fate transitions

systems biology

Orienting The Causal Relationship Between Imprecisely Measured Traits Using Genetic Instruments

Inference of the causal structure that induces correlations between two traits can be achieved by combining genetic associations with a mediation-based approach, as is done in the causal inference test (CIT) and others. However, we show that measurement error in the phenotypes can lead to mediation-based approaches inferring the wrong causal direction, and that increasing sample sizes has the adverse effect of increasing confidence in the wrong answer. Here we introduce an extension to Mendelian randomisation, a method that uses genetic associations in an instrumentation framework, that enables inference of the causal direction between traits, with some advantages. First, it is less susceptible to bias in the presence of measurement error; second, it is more statistically efficient; third, it can be performed using only summary level data from genome-wide association studies; and fourth, its sensitivity to measurement error can be evaluated. We apply the method to infer the causal direction between DNA methylation and gene expression levels. Our results demonstrate that, in general, DNA methylation is more likely to be the causal factor, but this result is highly susceptible to bias induced by systematic differences in measurement error between the platforms. We emphasise that, where possible, implementing MR and appropriate sensitivity analyses alongside other approaches such as CIT is important to triangulate reliable conclusions about causality.

systems biology

Evaluation and Design of Genome-wide CRISPR/Cas9 Knockout Screens

The adaptation of CRISPR/Cas9 technology to mammalian cell lines is transforming the study of human functional genomics. Pooled libraries of CRISPR guide RNAs (gRNAs), targeting human protein-coding genes and encoded in viral vectors, have been used to systematically create gene knockouts in a variety of human cancer and immortalized cell lines, in an effort to identify whether these knockouts cause cellular fitness defects. Previous work has shown that CRISPR screens are more sensitive and specific than pooled library shRNA screens in similar assays, but currently there exists significant variability across CRISPR library designs and experimental protocols. In this study, we re-analyze 17 genome-scale knockout screens in human cell lines from three research groups using three different genome-scale gRNA libraries, using the Bayesian Analysis of Gene Essentiality (BAGEL) algorithm to identify essential genes, to refine and expand our previously defined set of human core essential genes, from 360 to 684 genes. We use this expanded set of reference Core Essential Genes (CEG2), plus empirical data from six CRISPR knockout screens, to guide the design of a sequence-optimized gRNA library, the Toronto KnockOut version 3.0 (TKOv3) library. We demonstrate the high effectiveness of the library relative to reference sets of essential and nonessential genes as well as other screens using similar approaches. The optimized TKOv3 library, combined with the CEG2 reference set, provide an efficient, highly optimized platform for performing and assessing gene knockout screens in human cell lines.

systems biology

An evolutionary module in central metabolism

Metabolic enzyme function and evolution is influenced by the larger context of a biochemical pathway - deleterious mutations or perturbations in one enzyme can often be compensated by mutations to others. To explore strategies for mapping adaptive dependencies between enzymes, we used a combination of comparative genomics and experiments to examine interactions with the model metabolic enzyme Dihydrofolate Reductase (DHFR). Biochemically, DHFR shares a metabolic intermediate with numerous folate metabolic enzymes. In contrast, comparative genomics analyses of synteny and gene co-occurrence indicate a sparse pattern of evolutionary couplings in which DHFR is coupled to the enzyme thymidylate synthase (TYMS), but is relatively independent from the rest of folate metabolism. To test this apparent modularity, we used quantitative growth rate measurements and forward evolution in E. coli to demonstrate that the two enzymes are coupled to one another, and can adapt independently from the remainder of the genome. Mechanistically, the coupling between DHFR and TYMS is driven by a constraint wherein TYMS activity must not greatly exceed that of DHFR - both to avoid depletion of reduced folates and prevent accumulation of the metabolic intermediate dihydrofolate. Extending our comparative genomics analyses genome-wide reveals over 200 gene pairs with statistical signatures similar to DHFR/TYMS, suggesting the possibility that cellular pathways might be decomposed into small adaptive units.

systems biology

A personalized, multi-omics approach identifies genes involved in cardiac hypertrophy and heart failure

Identifying genes underlying complex diseases remains a major challenge. Biomarkers are typically identified by comparing average levels of gene expression in populations of healthy and diseased individuals. However, genetic diversities may undermine the effort to uncover genes with significant but individual contribution to the spectrum of disease phenotypes within a population. Here we leverage the Hybrid Mouse Diversity Panel (HMDP), a model system of 100+ genetically diverse strains of mice exhibiting different complex disease traits, to develop a personalized differential gene expression analysis that is able to identify disease-associated genes missed by traditional population-wide methods. The population-level and personalized approaches are compared for isoproterenol(ISO)-induced cardiac hypertrophy and heart failure using pre- and post-ISO gene expression and phenotypic data. The personalized approach identifies 36 Fold-Change (FC) genes predictive of the severity of cardiac hypertrophy, and enriched in genes previously associated with cardiac diseases in human. Strikingly, these genes are either up- or down-regulated at the individual strain level, and are therefore missed when averaging at the population level. Using insights from the gene regulatory network and protein-protein interactome, we identify Hes1 as a strong candidate FC gene. We validate its role by showing that even a mild knockdown of 20-40% of Hes1 can induce a dramatic reduction of hypertrophy by 80-90% in rat neonatal cardiac cells. These findings emphasize the importance of a personalized approach to identify causal genes underlying complex diseases as well as to develop personalized therapies.\n\nSignificanceA traditional approach to investigate the genetic basis of complex diseases is to look for genes with a global change in expression between diseased and healthy individuals. Here, we investigate individual changes of gene expression by inducing heart failure in 100 strains of genetically distinct mice. We find that genes associated to the severity of the disease are either up- or down-regulated across individuals and are therefore missed by a traditional population level approach. However, they are enriched in human cardiac disease genes and form a coregulated module strongly interacting with a cardiac hypertrophic signaling network in the human interactome. Our analysis demonstrates that individualized approaches are crucial to reveal all genes involved in the development of complex diseases.

systems biology