Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Layering genetic circuits to build a single cell, bacterial half adder

Gene regulation in biological systems is impacted by the cellular and genetic context-dependent effects of the biological parts which comprise the circuit. Here, we have sought to elucidate the limitations of engineering biology from an architectural point of view, with the aim of compiling a set of engineering solutions for overcoming failure modes during the development of complex, synthetic genetic circuits. Using a synthetic biology approach that is supported by computational modelling and rigorous characterisation, AND, OR and NOT biological logic gates were layered in both parallel and serial arrangements to generate a repertoire of Boolean operations that include NIMPLY, XOR, half adder and half subtractor logics in single cell. Subsequent evaluation of these near-digital biological systems revealed critical design pitfalls that triggered genetic context dependent effects, including 5 UTR interference and uncontrolled switch-on behaviour of {sigma}54 promoter. Importantly, this work provides a representative case study to the debugging of genetic context dependent effects through principles elucidated herein, thereby providing a rational design framework to program single prokaryotic cell with diversified digital operations.

Synthetic Biology

Predicting Causal Relationships from Biological Data: Applying Automated Casual Discovery on Mass Cytometry Data of Human Immune Cells

Learning the causal relationships that define a molecular system allows us to predict how the system will respond to different interventions. Distinguishing causality from mere association typically requires randomized experiments. Methods for automated causal discovery from limited experiments exist, but have so far rarely been tested in systems biology applications. In this work, we apply state-of-the art causal discovery methods on a large collection of public mass cytometry data sets, measuring intra-cellular signaling proteins of the human immune system and their response to several perturbations. We show how different experimental conditions can be used to facilitate causal discovery, and apply two fundamental methods that produce context-specific causal predictions. Causal predictions were reproducible across independent data sets from two different studies, but often disagree with the KEGG pathway databases. Within this context, we discuss the caveats we need to overcome for automated causal discovery to become a part of the routine data analysis in systems biology.

bioinformatics

Gene expression models based on a reference laboratory strain are bad predictors of Mycobacterium tuberculosis complex transcriptional diversity.

Species of the Mycobacterium tuberculosis complex (MTBC) kill more people every year than any other infectious disease. As a consequence of its global distribution and parallel evolution with the human host the bacteria is not genetically homogeneous. The observed genetic heterogeneity has relevance at different phenotypic levels, from gene expression to epidemiological dynamics. However current systems biology datasets have focused in the laboratory reference strain H37Rv. By using large expression datasets testing the role of almost two hundred transcription factors, we have constructed computational models to grab the expression dynamics of Mycobacterium tuberculosis H37Rv genes. However, we have found that many of those transcription factors are deleted or likely dysfunctional across strains of the MTBC. In accordance, we failed to predict expression changes in strains with a different genetic background when compared with experimental data. The results highlight the importance of designing systems biology approaches that take into account the tubercle bacilli, or any other pathogen, genetic diversity if we want to identify universal targets for vaccines, diagnostics and treatments.

systems biology

Advances in the integration of transcriptional regulatory information into genome-scale metabolic models

A major goal of systems biology is to build predictive computational models of cellular metabolism. Availability of complete genome sequences and wealth of legacy biochemical information has led to the reconstruction of genome-scale metabolic networks in the last 15 years for several organisms across the three domains of life. Due to paucity of information on kinetic parameters associated with metabolic reactions, the constraint-based modelling approach, flux balance analysis (FBA), has proved to be a vital alternative to investigate the capabilities of reconstructed metabolic networks. In parallel, advent of high-throughput technologies has led to the generation of massive amounts of omics data on transcriptional regulation comprising mRNA transcript levels and genome-wide binding profile of transcriptional regulators. A frontier area in metabolic systems biology has been the development of methods to integrate the available transcriptional regulatory information into constraint-based models of reconstructed metabolic networks in order to increase the predictive capabilities of computational models and understand the regulation of cellular metabolism. Here, we review the existing methods to integrate transcriptional regulatory information into constraint-based models of metabolic networks.

Systems Biology

Prenatal Bisphenol A Exposure in Mice Induces Multi-tissue Multi-omics Disruptions Linking to Cardiometabolic Disorders

The health impacts of endocrine disrupting chemicals (EDCs) remain debated and their tissue and molecular targets are poorly understood. Here we leveraged systems biology approaches to assess the target tissues, molecular pathways, and gene regulatory networks associated with prenatal exposure to the model EDC Bisphenol A (BPA). Prenatal BPA exposure led to scores of transcriptomic and methylomic alterations in the adipose, hypothalamus, and liver tissues in mouse offspring, with cross-tissue perturbations in lipid metabolism as well as tissue-specific alterations in histone subunits, glucose metabolism and extracellular matrix. Network modeling prioritized main molecular targets of BPA, including Pparg, Hnf4a, Esr1, and Fasn. Lastly, integrative analyses identified the association of BPA molecular signatures with cardiometabolic phenotypes in mouse and human. Our multi-tissue, multi-omics investigation provides strong evidence that BPA perturbs diverse molecular networks in central and peripheral tissues, and offers insights into the molecular targets that link BPA to human cardiometabolic disorders.\n\nAuthor summaryThe inability to pinpoint the mechanistic underpinnings of environmentally-induced diseases likely stems from the pleiotropic effects of chemicals such as BPA on diverse tissues and molecular space (transcriptome, epigenome, etc.). This makes it challenging to fully dissect their health impact and merits a call for modern big data approaches to examine environmental factors. Our data-driven study is the first unbiased, multi-tissue multiomic systems biology investigation of the molecular circuitry and mechanisms underlying offspring response to prenatal BPA exposure. Importantly, the incorporation of network-based modeling allows us to capture novel players in the regulation of BPA activities in vivo, and the integration with human disease association datasets helps bridge the molecular pathways affected by BPA with diverse human diseases. In doing so, our study provides compelling molecular evidence that developmental BPA exposure significantly perturbs metabolic and endocrine systems in the offspring, and supports BPA as one of the environmental factors involved in the developmental origins of health and disease (DOHaD).

systems biology

The feasibility and stability of large complex biological networks: a random matrix approach

In his theoretical work of the 70s, Robert May introduced a Random Matrix Theory (RMT) approach for studying the stability of large complex biological systems. Unlike the established paradigm, May demonstrated that complexity leads to instability in generic models of biological networks. The RMT approach has since similarly been applied in many disciplines. Central to the approach is the famous \"circular law\" that describes the eigenvalue distribution of an important class of random matrices. However the \"circular law\" generally does not apply for ecological and biological systems in which density-dependence (DD) operates. Here we directly determine the far more complicated eigenvalue distributions of complex DD systems. A simple mathematical solution falls out, that allows us to explore the connection between feasible systems (i.e., having all equilibrium populations positive) and stability. In particular, for these RMT systems, almost all feasible systems are stable. The degree of stability, or resilience, is shown to depend on the minimum equilibrium population, and not directly on factors such as network topology.

ecology

Synthesis and patterning of tunable multiscale materials with engineered cells

A major challenge in materials science is to create self-assembling, functional, and environmentally responsive materials which can be patterned across multiple length scales. Natural biological systems, such as biofilms, shells, and skeletal tissues, implement dynamic regulatory programs to assemble complex multiscale materials comprised of living and non-living components1-9. Such systems can provide inspiration for the design of heterogeneous functional systems which integrate biotic and abiotic materials via hierarchical self-assembly. Here, we present a synthetic-biology platform for synthesizing and patterning self-assembled functional amyloid materials across multiple length scales with bacterial biofilms. We engineered Escherichia coli curli amyloid production under the tight control of synthetic regulatory circuits and interfaced amyloids with inorganic materials to create a biofilm-based electrical switch whose conductance can be selectively toggled by specific environmental signals. Furthermore, we externally tuned synthetic biofilms to build nanoscale amyloid biomaterials with different structure and composition through the controlled expression of their constituent subunits with artificial gene circuits. By using synthetic cell-cell communication, our engineered biofilms can also autonomously manufacture dynamic materials whose structure and composition change with time. In addition, we show that by combining subunit-level protein engineering, controlled genetic expression of self-assembling subunit proteins, and macroscale spatial gradients, synthetic biofilms can pattern protein biomaterials across multiple length scales. This work lays a foundation for synthesizing, patterning, and controlling composite materials with engineered biological systems. We envision that this approach can be expanded to other cellular and biomaterials contexts for the construction of self-organizing, environmentally responsive, and tunable multiscale composite materials with heterogeneous functionalities.

Synthetic Biology

BioLogic, a parallel approach to cell-based logic gates

AbstractIn vivo logic gates have proven difficult to combine into larger devices. Our cell-based logic system, BioLogic, decomposes a large circuit into a collection of small subcircuits working in parallel, each subcircuit responding to a different combination of inputs. A final global output is then generated by a combination of the responses. Using BioLogic, for the first time a completely functional 3-bit full adder and full subtractor were generated using Escherichia coli cells; as well as a calculator-style display that shows a numeric result, from 0 to 7, when the proper 3 bit binary inputs are introduced into the system. BioLogic demonstrates the use of a parallel approach for the design of cell-based logic gates that facilitates the generation and analysis of complex processes, without the need for complex genetic engineering.

synthetic biology

Intrinsic limitations in mainstream methods of identifying network motifs in biology

Network motifs are connectivity structures that occur with significantly higher frequency than chance, and are thought to play important roles in complex biological networks, for example in gene regulation, interactomes, and metabolomes. Network motifs may also become pivotal in the rational design and engineering of complex biological systems underpinning the field of synthetic biology. Distinguishing true motifs from arbitrary substructures, however, remains a challenge. Here we demonstrate both theoretically and empirically that implicit assumptions present in mainstream methods for motif identification do not necessarily hold, with the ramification that motif studies using these mainstream methods are less able to effectively differentiate between spurious results and events of true statistical significance than is often presented. We show that these difficulties cannot be overcome without revising the methods of statistical analysis used to identify motifs. The implications of these findings are therefore far-reaching across diverse areas of biology.

systems biology

Experimental evolution of Escherichia coli harboring an ancient translation protein

The ability to design synthetic genes and engineer biological systems at the genome scale opens new means by which to characterize phenotypic states and the responses of biological systems to perturbations. One emerging method involves inserting artificial genes into bacterial genomes, and examining how the genome and its new genes adapt to each other. Here we report the development and implementation of a modified approach to this method, in which phylogenetically inferred genes are inserted into a microbial genome, and laboratory evolution is then used to examine the adaptive potential of the resulting hybrid genome. Specifically, we engineered an approximately 700-million-year old inferred ancestral variant of tufB, an essential gene encoding Elongation Factor Tu, and inserted it in a modern Escherichia coli genome in place of the native tufB gene. While the ancient homolog was not lethal to the cell, it did cause a two-fold decrease in organismal fitness, mainly due to reduced protein dosage. We subsequently evolved replicate hybrid bacterial populations for 2,000 generations in the laboratory, and examined the adaptive response via fitness assays, whole-genome sequencing, proteomics and biochemical assays. Hybrid lineages exhibit a general adaptive strategy in which the fitness cost of the ancient gene was ameliorated in part by up-regulation of protein production. We expect that this ancient-modern recombinant method may pave the way for the synthesis of organisms that exhibit ancient phenotypes, and that laboratory evolution of these organisms may prove useful in elucidating insights into historical adaptive processes.

Evolutionary Biology

Quantitative Modeling of Integrase Dynamics Using a Novel Python Toolbox for Parameter Inference in Synthetic Biology

In systems and synthetic biology, it is common to build chemical reaction network (CRN) models of biochemical circuits and networks. Although automation and other high-throughput techniques have led to an abundance of data enabling data-driven quantitative modeling and parameter estimation, the intense amount of simulation needed for these methods still frequently results in a computational bottleneck. Here we present bioscrape (Bio-circuit Stochastic Single-cell Reaction Analysis and Parameter Estimation) - a Python package for fast and flexible modeling and simulation of highly customizable chemical reaction networks. Specifically, bioscrape supports deterministic and stochastic simulations, which can incorporate delay, cell growth, and cell division. All functionalities - reaction models, simulation algorithms, cell growth models, partioning models, and Bayesian inference - are implemented as interfaces in an easily extensible and modular object-oriented framework. Models can be constructed via Systems Biology Markup Language (SBML) or specified programmatically via a Python API. Simulation run times obtained with the package are comparable to those obtained using C code - this is particularly advantageous for computationally expensive applications such as Bayesian inference or simulation of cell lineages. We first show the packages simulation capabilities on a variety of example simulations of stochastic gene expression. We then further demonstrate the package by using it to do parameter inference on a model of integrase enzyme-mediated DNA recombination dynamics with experimental data. The bioscrape package is publicly available online (https://github.com/biocircuits/bioscrape) along with more detailed documentation and examples.

synthetic biology

High sensitivity quantitative proteomics using accumulated ion monitoring and automated multidimensional nano-flow chromatography

Quantitative proteomics using high-resolution and accuracy mass spectrometry promises to transform our understanding of biological systems and disease. Recent development of parallel reaction monitoring (PRM) using hybrid instruments substantially improved the specificity of targeted mass spectrometry. Combined with high-efficiency ion trapping, this approach also provided significant improvements in sensitivity. Here, we investigated the effects of ion isolation and accumulation on the sensitivity and quantitative accuracy of targeted proteomics using the recently developed hybrid quadrupole-Orbitrap-linear ion trap mass spectrometer. We leveraged ultra-high efficiency nano-electrospray ionization under optimized conditions to achieve yoctomolar sensitivity with more than seven orders of linear quantitative accuracy. To enable sensitive and specific targeted mass spectrometry, we implemented an automated, scalable two-dimensional (2D) ion exchange-reversed phase nano-scale chromatography system. We found that 2D chromatography improved the sensitivity and accuracy of both PRM and an intact precursor scanning mass spectrometry method, termed accumulated ion monitoring (AIM), by more than 100-fold. Combined with automated 2D nano-scale chromatography, AIM achieved sub-attomolar limits of detection of endogenous proteins in complex biological proteomes. This allowed quantitation of absolute abundance of the human transcription factor MEF2C at approximately 100 molecules/cell, and determination of its phosphorylation stoichiometry from as little as 1 g of extracts isolated from 10,000 human cells. The combination of automated multidimensional nano-scale chromatography and targeted mass spectrometry should enable ultra-sensitive high-accuracy quantitative proteomics of complex biological systems and diseases.

molecular biology

Insights into the relation between noise and biological complexity

Understanding under which conditions the increase of systems complexity is evolutionary advantageous, and how this trend is related to the modulation of the intrinsic noise, are fascinating issues of utmost importance for synthetic and systems biology. To get insights into these matters, we analyzed chemical reaction networks with different topologies and degrees of complexity, interacting or not with the environment. We showed that the global level of fluctuations at the steady state, as measured by the sum of the Fano factors of the number of molecules of all species, is directly related to the topology of the network. For systems with zero deficiency, this sum is constant and equal to the rank of the network. For higher deficiencies, we observed an increase or decrease of the fluctuation levels according to the values of the reaction fluxes that link internal species, multiplied by the associated stoichiometry. We showed that the noise is reduced when the fluxes all flow towards the species of higher complexity, whereas it is amplified when the fluxes are directed towards lower complexity species.\n\nPACS numbers: 02.50.Ey, 05.10.Gg, 05.40.Ca, 87.18.-h

systems biology

Computational modelling of atherosclerosis: developing a community resource

RationaleAtherosclerosis is a dynamical process that emerges from the interplay between lipid metabolism, inflammation and innate immunity. The arterial location of atherosclerosis makes it logistically and ethically difficult to study in vivo. To improve our understanding of the disease, we must find alternative ways to investigate its progression. There is currently no computational model of atherosclerosis openly available to the research community for use in future studies and for refinement and development.\n\nObjectiveHere we develop the first predictive computational model to be made openly available and demonstrate its use for therapeutic hypothesis generation.\n\nMethods and ResultsWe compiled a dataset of relevant interactions from the literature along with available parameters. These were used to build a network model describing atherosclerotic plaque development. A visual map of the network model was produced using the Systems Biology Graphical Notation (SBGN) and a dynamic mathematical description of the network model that enables us to simulate plaque growth was developed and is made available using the Systems Biology Markup Language (SBML). We used this model to investigate whether multi-drug therapeutic interventions could be identified that stimulate plaque regression. The model produced comprised 20 cell types and 41 proteins with 89 species in total. The visual map is available for reuse and refinement using the SBGN Markup Language standard format and the mathematical model is available using the SBML standard format. We used a genetic algorithm to identify a multi-drug intervention hypothesis comprising five drugs that comprehensively reverse plaque growth within the model.\n\nConclusionsWe have produced the first predictive mathematical and computational model of atherosclerosis that can be reused and refined by the cardiovascular research community. We demonstrated its potential as a tool for future studies of cardiovascular disease by using it to identify multi-drug intervention hypotheses.\n\nSubject CodesAtherosclerosis, Computational Biology, Lipids and Cholesterol, Cell Signaling/Signal Transduction, Cardiovascular Disease

physiology

Dynamic Network Analysis of the 4D Nucleome

MotivationFor many biological systems, it is essential to capture simultaneously the function, structure, and dynamics in order to form a comprehensive understanding of underlying phenomena. The dynamical interaction between 3D genome spatial structure and transcriptional activity creates a genomic signature that we refer to as the four-dimensional organization of the nucleus, or 4D Nucleome (4DN). The study of 4DN requires assessment of genome-wide structure and gene expression as well as development of new approaches for data analysis.\n\nResultsWe propose a dynamic multilayer network approach to study the co-evolution of form and function in the 4D Nucleome. We model the dynamic biological system as a temporal network with node dynamics, where the network topology is captured by chromosome conformation (Hi-C), and the function of a node is measured by RNA sequencing (RNA-seq). Network-based approaches such as von Neumann graph entropy, network centrality, and multilayer network theory are applied to reveal universal patterns of the dynamic genome. Our model integrates knowledge of genome structure and gene expression along with temporal evolution and leads to a description of genome behavior on a system wide level. We illustrate the benefits of our model via a real biological dataset on MYOD1-mediated reprogramming of human fibroblasts into the myogenic lineage. We show that our methods enable better predictions on form-function relationships and refine our understanding on how cell dynamics change during cellular reprogramming.\n\nAvailability: The software is available upon request.\n\nContactindikar@umich.edu\n\nSupplementary informationSee Supplementary Material.

bioinformatics

Identifying protein subsets and features responsible for improved drug repurposing accuracies using the CANDO platform

Drug repurposing is a valuable tool for combating the slowing rates of novel therapeutic discovery. The Computational Analysis of Novel Drug Opportunities (CANDO) platform performs shotgun repurposing of 3,733 drugs/compounds that map to 2,030 indications/diseases by predicting their interactions with 46,784 protein structures and relating them via proteomic interaction signatures. The accuracy of the CANDO platform is evaluated using our benchmarking protocol that assesses indication accuracies based on whether or not pairs of drugs associated with the same indication can be captured within a certain cutoff, which is a measure of the drug repurposing recovery rate. To identify subsets of proteins that exhibit the same therapeutic effectiveness as the full set, groups of 8 proteins were randomly selected and subsequently benchmarked 50 times. The resulting protein sets were ranked according to average indication accuracy, pairwise accuracy, and coverage (count of indications with non-zero accuracy). The best 50 subsets of 8 according to each metric were progressively combined into supersets after each iteration and benchmarked. These supersets yield up to 14% improvement in benchmarking accuracy, and represent a 100-1,000 fold reduction in the number of proteins relative to the full set. Protein supersets optimized using independent compound libraries derived from the full library were cross-tested and were shown to reproduce the performance relative to using all 46,784 proteins, indicating that these reduced size supersets are broadly applicable for characterizing drug behavior. Further analysis revealed that sets comprised of proteins with more equitably diverse ligand interactions are important for describing drug behavior. Our work elucidates the role of particular protein subsets and corresponding ligand interactions that play a role in computational drug repurposing, and paves the way for the use of machine learning approaches to further improve the accuracy of the CANDO platform and its repurposing potential.\n\nAuthor summaryDrug repurposing is a valuable approach for ameliorating the current problems plaguing drug discovery. We introduce a novel protein subset analysis pipeline that allows us to elucidate features important for drug repurposing accuracies using the Computational Analysis of Novel Drug Opportunities (CANDO) platform. Our platform relates drugs based on the similarity of their interactions with a diverse library of proteins. We subjected all proteins in the platform to a splitting and ranking protocol that ranked protein subsets based on their benchmarking performance. Further analysis of the best performing protein subsets revealed that the most useful proteins for describing how small molecule compounds behave in biological systems are those that are predicted to interact with a structurally diverse range of ligands. We hypothesize that this is a consequence of the multitarget nature of drugs and, conversely, the implied promiscuity of proteins in biological systems. These results may be used to make drug discovery more accurate and efficient by alleviating some of its bottlenecks, bringing us one step further in better understanding how drugs behave in the context of their environments.

bioinformatics

SuperDCA for genome-wide epistasis analysis

The potential for genome-wide modeling of epistasis has recently surfaced given the possibility of sequencing densely sampled populations and the emerging families of statistical interaction models. Direct coupling analysis (DCA) has earlier been shown to yield valuable predictions for single protein structures, and has recently been extended to genome-wide analysis of bacteria, identifying novel interactions in the co-evolution between resistance, virulence and core genome elements. However, earlier computational DCA methods have not been scalable to enable model fitting simultaneously to 104-105 polymorphisms, representing the amount of core genomic variation observed in analyses of many bacterial species. Here we introduce a novel inference method (SuperDCA) which employs a new scoring principle, efficient parallelization, optimization and filtering on phylogenetic information to achieve scalability for up to 105 polymorphisms. Using two large population samples of Streptococcus pneumoniae, we demonstrate the ability of SuperDCA to make additional significant biological findings about this major human pathogen. We also show that our method can uncover signals of selection that are not detectable by genome-wide association analysis, even though our analysis does not require phenotypic measurements. SuperDCA thus holds considerable potential in building understanding about numerous organisms at a systems biological level.\n\nAuthor SummaryRecent work has demonstrated the emerging potential in statistical genome-wide modeling to uncover co-selection and epistatic interactions between polymorphisms in bacterial chromosomes from densely sampled population data. Here we develop the Potts model based approach further into a fully mature computational method which can be applied to most existing bacterial population genomic data sets in a straightforward manner. Our advances are relying on more efficient parameter scoring, highly optimized and parallelized open source C++ code, which does not rely on the computation-intensive polymorphism subsampling approximations used earlier. By analyzing the two largest available population samples of Streptococcus pneumoniae (the pneumococcus), we highlight several biological discoveries related to the survival of the pneumococcus and co-evolution of penicillin-binding loci, which were not uncovered by the earlier analyses. Our method holds considerable potential for building understanding about numerous organisms at a systems biological level.

genomics

Integrated deep learned transcriptomic and structure-based predictor of clinical trials outcomes

Despite many recent advances in systems biology and a marked increase in the availability of high-throughput biological data, the productivity of research and development in the pharmaceutical industry is on the decline. This is primarily due to clinical trial failure rates reaching up to 95% in oncology and other disease areas. We have developed a comprehensive analytical and computational pipeline utilizing deep learning techniques and novel systems biology analytical tools to predict the outcomes of phase I/II clinical trials. The pipeline predicts the side effects of a drug using deep neural networks and estimates drug-induced pathway activation. It then uses the predicted side effect probabilities and pathway activation scores as an input to train a classifier which predicts clinical trial outcomes. This classifier was trained on 577 transcriptomic datasets and has achieved a cross-validated accuracy of 0.83. When compared to a direct gene-based classifier, our multi-stage approach dramatically improves the accuracy of the predictions. The classifier was applied to a set of compounds currently present in the pipelines of several major pharmaceutical companies to highlight potential risks in their portfolios and estimate the fraction of clinical trials that were likely to fail in phase I and II.

bioinformatics