Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Complex-centric proteome profiling by SEC-SWATH-MS

Proteins are major effectors and regulators of biological processes that can elicit multiple functions depending on their interaction with other proteins. The organization of proteins into macromolecular complexes and their quantitative distribution across these complexes is, therefore, of great biological and clinical significance.\n\nIn this paper we describe an integrated experimental and computational technique to quantify hundreds of protein complexes in a single operation. The method consists of size exclusion chromatography (SEC) to fractionate native protein complexes, SWATH/DIA mass spectrometry to precisely quantify the proteins in each SEC fraction and the computational framework CCprofiler to detect and quantify protein complexes by error-controlled, complex-centric analysis using prior information from generic protein interaction maps.\n\nOur analysis of the HEK293 cell line proteome delineates 462 complexes composed of 2127 protein subunits. The technique identifies novel subcomplexes and assembly intermediates of central regulatory complexes while assessing the quantitative subunit distribution across them. We make the toolset CCprofiler freely accessible, and provide a web platform, SECexplorer, for custom exploration of the HEK293 proteome modularity.\n\nHighlightsO_LIIntroduction of the concept of complex-centric proteome profiling\nC_LIO_LIDevelopment of CCprofiler, a software framework for complex-centric data analysis\nC_LIO_LIDetection and quantification of subunit distribution of 462 distinct protein complexes containing 2127 proteins from a SEC-SWATH-MS dataset of HEK293 cells, and identification of novel complex variants such as assembly intermediates\nC_LIO_LIStatistical target-decoy model to estimate accurate false discovery rates for complexes quantified by complex-centric analysis\nC_LIO_LISECexplorer, an online platform to support custom complex-centric exploration of SEC-SWATH-MS datasets.\nC_LI

systems biology

Accurate, Sensitive, and Precise Multiplexed Proteomics using the Complement Reporter Ion Cluster

Quantitative analysis of proteomes across multiple time points, organelles, and perturbations is essential for understanding both fundamental biology and disease states. The development of isobaric tags (e.g. TMT) have enabled the simultaneous measurement of peptide abundances across several different conditions. These multiplexed approaches are promising in principle because of advantages in throughput and measurement quality. However, in practice existing multiplexing approaches suffer from key limitations. In its simple implementation (TMT-MS2), measurements are distorted by chemical noise leading to poor measurement accuracy. The current state-of- the-art (TMT-MS3) addresses this, but requires specialized quadrupole-iontrap-Orbitrap instrumentation. The complement reporter ion approach (TMTc) produces high accuracy measurements and is compatible with many more instruments, like quadrupole-Orbitraps. However, the required deconvolution of the TMTc cluster leads to poor measurement precision. Here, we introduce TMTc+, which adds the modeling of the MS2- isolation step into the deconvolution algorithm. The resulting measurements are comparable in precision to TMT-MS3/MS2. The improved duty cycle, and lower filtering requirements make TMTc+ more sensitive than TMT-MS3 and comparable with TMT-MS2. At the same time, unlike TMT-MS2, TMTc+ is exquisitely able to distinguish signal from chemical noise even outperforming TMT-MS3. Lastly, we compare TMTc+ to quantitative label-free proteomics of total HeLa lysate and find that TMTc+ quantifies 7.8k versus 3.9k proteins in a 5-plex sample. At the same time the median coefficient of variation improves from 13% to 4%. Thus, TMTc+ advances quantitative proteomics by enabling accurate, sensitive, and precise multiplexed experiments on more commonly used instruments.

systems biology

Random Walk With Restart On Multiplex And Heterogeneous Biological Networks

Recent years have witnessed an exponential growth in the number of identified interactions between biological molecules. These interactions are usually represented as large and complex networks, calling for the development of appropriated tools to exploit the functional information they contain. Random walk with restart is the state-of-the-art guilt-by-association approach. It explores the network vicinity of gene/protein seeds to study their functions, based on the premise that nodes related to similar functions tend to lie close to each others in the networks.\n\nIn the present study, we extended the random walk with restart algorithm to multiplex and heterogeneous networks. The walk can now explore different layers of physical and functional interactions between genes and proteins, such as protein-protein interactions and co-expression associations. In addition, the walk can also jump to a network containing different sets of edges and nodes, such as phenotype similarities between diseases.\n\nWe devised a leave-one-out cross-validation strategy to evaluate the algorithms abilities to predict disease-associated genes. We demonstrate the increased performances of the multiplex-heterogeneous random walk with restart as compared to several random walks on monoplex or heterogeneous networks. Overall, our framework is able to leverage the different interaction sources to outperform current approaches.\n\nFinally, we applied the algorithm to predict genes candidate for being involved in the Wiedemann-Rautenstrauch syndrome, and to explore the network vicinity of the SHORT syndrome.\n\nThe source code and the software are freely available at: https://github.com/alberto-valdeolivas/RWR-MH.

systems biology

The Soft Vertex Classification for Active Module Identification Problem

MotivationIntegrative network methods are commonly used for interpretation of high-throughput experimental biological data: transcriptomics, proteomics, metabolomics and others. One of the common approaches consists in finding a connected subnetwork of a global interaction network that best encompasses significant individual changes in the data and represents a so-called active module. Usually methods implementing this approach find a single subnetwork and thus solve a hard classification problem for vertices. This subnetwork inherently contains erroneous vertices, while no instrument is provided to estimate the confidence level of any particular vertex inclusion. To address this issue, in the current study we consider the active module problem as a soft classification problem. We propose a method to estimate probabilities of each vertex to belong to the active module based on Markov chain Monte Carlo subnetwork sampling.\n\nResultsThe proposed method allows to estimate the probability that an individual vertex belongs to the active module as well as the false discovery rate (FDR) for a given set of vertices. Given the estimated probabilities, it becomes possible to provide a connected subgraph in a consistent manner for any given FDR level: no vertex can disappear when the FDR level is relaxed. We show on simulated dataset that the proposed method has good computational performance and high classification accuracy. As an example of the performance of our method on real data, we run it on a protein-protein interaction network together with a gene expression DLBCL dataset. The results are consistent with the previous studies while, at the same time, the proposed approach is more flexible. Source code is available at https://github.com/ctlab/mcmcRanking under MIT licence.

systems biology

Sex-biased long non-coding RNAs negatively correlated with sex-opposite protein coding gene co-expression networks in Diversity Outbred mouse liver

Sex differences in liver gene expression and disease susceptibility are regulated by pituitary growth hormone secretion patterns, which activate sex-dependent liver transcription factors and establish sex-specific chromatin states. Ablation of pituitary hormone by hypophysectomy (hypox) has identified two major classes of liver sex-biased genes, defined by their sex-dependent positive or negative responses to hypox, respectively; however, the mechanisms that determine the hypox responsiveness of each gene class are unknown. Here, we sought to discover candidate regulatory long noncoding RNAs (lncRNAs) that control hypox responsiveness. First, we used mouse liver RNA-seq data for 30 different biological conditions to discover gene structures and expression patterns for ~15,500 liver-expressed lncRNAs, including antisense and intragenic lncRNAs, as well as lncRNAs that overlap active enhancers, marked by enhancer RNAs. We identified >200 robust sex-specific liver lncRNAs, including 157 whose expression is regulated during postnatal liver development or is subject to circadian oscillations. Next, we utilized the high natural allelic variance of Diversity Outbred (DO) mice, a multi-parental outbred population, to discover tightly co-expressed clusters of sex-specific protein-coding genes (gene modules) in male liver, and separately, in female liver. Sex differences in the gene modules identified were extensive. Remarkably, many gene modules were strongly enriched for male-specific or female-specific genes belonging to a single hypox-response classes, indicating that the genetic heterogeneity of DO mice captures responsiveness to hypox. Hypox-responsiveness was shown to be facilitated by multiple, distinct gene regulatory mechanisms, indicating its complex nature. Further, we identified 16 sex-specific lncRNAs whose expression across DO mouse livers showed an unexpected significant negative correlation with protein-coding gene modules enriched for genes of the opposite-sex bias and inverse hypox response class, indicating strong negative regulatory potential for these lncRNAs. Thus, we used a genetically diverse outbred mouse population to discover tightly co-expressed sex-specific gene modules that reveal broad characteristics of gene regulation related to responsiveness to hypox, and generated testable hypotheses for regulatory roles of sex-biased liver lncRNAs that control the sex-bias in liver gene expression.

systems biology

Characterization of missing values in untargeted MS-based metabolomics data and evaluation of missing data handling strategies

BACKGROUNDUntargeted mass spectrometry (MS)-based metabolomics data often contain missing values that reduce statistical power and can introduce bias in epidemiological studies. However, a systematic assessment of the various sources of missing values and strategies to handle these data has received little attention. Missing data can occur systematically, e.g. from run day-dependent effects due to limits of detection (LOD); or it can be random as, for instance, a consequence of sample preparation.\n\nMETHODSWe investigated patterns of missing data in an MS-based metabolomics experiment of serum samples from the German KORA F4 cohort (n = 1750). We then evaluated 31 imputation methods in a simulation framework and biologically validated the results by applying all imputation approaches to real metabolomics data. We examined the ability of each method to reconstruct biochemical pathways from data-driven correlation networks, and the ability of the method to increase statistical power while preserving the strength of established genetically metabolic quantitative trait loci.\n\nRESULTSRun day-dependent LOD-based missing data accounts for most missing values in the metabolomics dataset. Although multiple imputation by chained equations (MICE) performed well in many scenarios, it is computationally and statistically challenging. K-nearest neighbors (KNN) imputation on observations with variable pre-selection showed robust performance across all evaluation schemes and is computationally more tractable.\n\nCONCLUSIONMissing data in untargeted MS-based metabolomics data occur for various reasons. Based on our results, we recommend that KNN-based imputation is performed on observations with variable pre-selection since it showed robust results in all evaluation schemes.\n\nKey messagesO_LIUntargeted MS-based metabolomics data show missing values due to both batch-specific LOD-based and non-LOD-based effects.\nC_LIO_LIStatistical evaluation of multiple imputation methods was conducted on both simulated and real datasets.\nC_LIO_LIBiological evaluation on real data assessed the ability of imputation methods to preserve statistical inference of biochemical pathways and correctly estimate effects of genetic variants on metabolite levels.\nC_LIO_LIKNN-based imputation on observations with variable pre-selection and K = 10 showed robust performance for all data scenarios across all evaluation schemes.\nC_LI

systems biology

The Strehler-Mildvan correlation is nothing but a fitting artifact

Gompertz empirical law of mortality is often used in practical research to parametrize survival fraction as a function of age with the help of just two quantities: the Initial Mortality Rate (IMR) and the Gompertz exponent, inversely proportional to the Mortality Rate Doubling Time (MRDT). The IMR is often found to be inversely related to the Gompertz exponent, which is the dependence commonly referred to as Strehler-Mildvan (SM) correlation. In this paper, we address fundamental uncertainties of the Gompertz parameters inference from experimental Kaplan-Meier plots and show, that a least squares fit often leads to an ill-defined non-linear optimization problem, which is extremely sensitive to sampling errors and the smallest systematic demographic variations. Therefore, an analysis of consequent repeats of the same experiments in the same biological conditions yields the whole degenerate manifold of possible Gompertz parameters. We find that whenever the average lifespan of species greatly exceeds MRDT, small random variations in the survival records produce large deviations in the identified Gompertz parameters along the line, corresponding to the set of all possible IMR and MRDT values, roughly compatible with the properly determined value of average lifespan in experiment. The best fit parameters in this case turn out to be related by a form of SM correlation. Therefore, we have to conclude that the combined property, such as the average lifespan in the group, rather than IMR and MRDT values separately, may often only be reliably determined via experiments, even in a perfectly homogeneous animal cohort due to its finite size and/or low age-sampling frequency, typical for modern high-throughput settings. We support our findings with careful analysis of experimental survival records obtained in cohorts of C. elegans of different sizes, in control groups and under the influence of experimental therapies or environmental conditions. We argue that since, SM correlation may show up as a consequence of the fitting degeneracy, its appearance is not limited to homogeneous cohorts. In fact, the problem persists even beyond the simple Gompertz mortality law. We show that the same degeneracy occurs exactly in the same way, if a more advanced Gompertz-Makeham aging model is employed to improve the modeling. We explain how SM type of relation between the demographic parameters may still be observed even in extremely large cohorts with immense statistical power, such as in human census datasets, provided that systematic historical changes are weak in nature and lead to a gradual change in the mean lifespan.

Systems Biology

Single-Cell Transcriptomic Analysis of Human Lung Reveals Complex Multicellular Changes During Pulmonary Fibrosis

Pulmonary fibrosis is a devastating disorder that results in the progressive replacement of normal lung tissue with fibrotic scar. Available therapies slow disease progression, but most patients go on to die or require lung transplantation. Single-cell RNA-seq is a powerful tool that can reveal cellular identity via analysis of the transcriptome, but its ability to provide biologically or clinically meaningful insights in a disease context is largely unexplored. Accordingly, we performed single-cell RNA-seq on lung tissue obtained from eight transplant donors and eight recipients with pulmonary fibrosis and one bronchoscopic cryobiospy sample. Integrated single-cell transcriptomic analysis of donors and patients with pulmonary fibrosis identified the emergence of distinct populations of epithelial cells and macrophages that were common to all patients with lung fibrosis. Analysis of transcripts in the Wnt pathway suggested that within the same cell type, Wnt secretion and response are restricted to distinct non-overlapping cells, which was confirmed using in situ RNA hybridization. Single-cell RNA-seq revealed heterogeneity within alveolar macrophages from individual patients, which was confirmed by immunohistochemistry. These results support the feasibility of discovery-based approaches applying next generation sequencing technologies to clinically obtained samples with a goal of developing personalized therapies.\n\nOne Sentence SummarySingle-cell RNA-seq applied to tissue from diseased and donor lungs and a living patient with pulmonary fibrosis identifies cell type-specific disease-associated molecular pathways.

systems biology

Evolution of resilience in protein interactomes across the tree of life

Phenotype robustness to environmental fluctuations is a common biological phenomenon. Although most phenotypes involve multiple proteins that interact with each other, the basic principles of how such interactome networks respond to environmental unpredictability and change during evolution are largely unknown. Here we study interactomes of 1,840 species across the tree of life involving a total of 8,762,166 protein-protein interactions. Our study focuses on the resilience of interactomes to network failures and finds that interactomes become more resilient during evolution, indicating that a species position in the tree of life is predictive of how robust its interactome is to network failures. In bacteria, we find that a more resilient interactome is in turn associated with the greater ability of the organism to survive in a more complex, variable and competitive environment. We find that at the protein family level, proteins exhibit a coordinated rewiring of interactions over time and that a resilient interactome arises through gradual change of the network topology. Our findings have implications for understanding molecular network structure both in the context of evolution and environment.\n\nSignificance StatementThe interactome network of protein-protein interactions captures the structure of molecular machinery that underlies organismal complexity. The resilience to network failures is a critical property of the interactome as the breakdown of interactions may lead to cell death or disease. By studying interactomes from 1,840 species across the tree of life, we find that evolution leads to more resilient interactomes, providing evidence for a longstanding hypothesis that interactomes evolve favoring robustness against network failures. We find that a highly resilient interactome has a beneficial impact on the organisms survival in complex, variable, and competitive habitats. Our findings reveal how interactomes change through evolution and how these changes affect their response to environmental unpredictability.

systems biology

Phenotype-Driven Transitions In Regulatory Network Structure

Complex traits and diseases like human height or cancer are often not caused by a single mutation or genetic variant, but instead arise from multiple factors that together functionally perturb the underlying molecular network. Biological networks are known to be highly modular and contain dense \"communities\" of genes that carry out cellular processes, but these structures change between tissues, during development, and in disease. While many methods exist for inferring networks, we lack robust methods for quantifying changes in network structure. Here, we describe ALPACA (ALtered Partitions Across Community Architectures), a method for comparing two genome-scale networks derived from different phenotypic states to identify condition-specific modules. In simulations, ALPACA leads to more nuanced, sensitive, and robust module discovery than currently available network comparison methods. We used ALPACA to compare transcriptional networks in three contexts: angiogenic and non-angiogenic subtypes of ovarian cancer, human fibroblasts expressing transforming viral oncogenes, and sexual dimorphism in human breast tissue. In each case, ALPACA identified modules enriched for processes relevant to the phenotype. For example, modules specific to angiogenic ovarian tumors were enriched for genes associated with blood vessel development, interferon signaling, and flavonoid biosynthesis. In comparing the modular structure of networks in female and male breast tissue, we found that female breast has distinct modules enriched for genes involved in estrogen receptor and ERK signaling. The functional relevance of these new modules indicate that not only does phenotypic change correlate with network structural changes, but also that ALPACA can identify such modules in complex networks.\n\nSignificance statementDistinct phenotypes are often thought of in terms of unique patterns of gene expression. But the expression levels of genes and proteins are driven by networks of interacting elements, and changes in expression are driven by changes in the structure of the associated networks. Because of the size and complexity of these networks, identifying functionally significant changes in network topology has been an ongoing challenge. We describe a new method for comparing networks derived from related conditions, such as healthy and disease tissue, and identifying emergent modules associated with the phenotypic differences between the conditions. We show that this method can find both known and previously unreported pathways involved in three contexts: ovarian cancer, tumor viruses, and breast tissue development.

systems biology

Differential Community Detection in Paired Biological Networks

MotivationBiological networks unravel the inherent structure of molecular interactions which can lead to discovery of driver genes and meaningful pathways especially in cancer context. Often due to gene mutations, the gene expression undergoes changes and the corresponding gene regulatory network sustains some amount of localized re-wiring. The ability to identify significant changes in the interaction patterns caused by the progression of the disease can lead to the revelation of novel relevant signatures.\n\nMethodsThe task of identifying differential sub-networks in paired biological networks (A:control,B:case) can be re-phrased as one of finding dense communities in a single noisy differential topological (DT) graph constructed by taking absolute difference between the topological graphs of A and B. In this paper, we propose a fast two-stage approach, namely Differential Community Detection (DCD), to identify differential sub-networks as differential communities in a de-noised version of the DT graph. In the first stage, we iteratively re-order the nodes of the DT graph to determine approximate block diagonals present in the DT adjacency matrix using neighbourhood information of the nodes and Jaccard similarity. In the second stage, the ordered DT adjacency matrix is traversed along the diagonal to remove all the edges associated with a node, if that node has no immediate edges within a window. We then apply community detection methods on this de-noised DT graph to discover differential sub-networks as communities.\n\nResultsOur proposed DCD approach can effectively locate differential sub-networks in several simulated paired random-geometric networks and various paired scale-free graphs with different power-law exponents. The DCD approach easily outperforms community detection methods applied on the original noisy DT graph and recent statistical techniques in simulation studies. We applied DCD method on two real datasets: a) Ovarian cancer dataset to discover differential DNA co-methylation sub-networks in patients and controls; b) Glioma cancer dataset to discover the difference between the regulatory networks of IDH-mutant and IDH-wild-type. We demonstrate the potential benefits of DCD for finding network-inferred bio-markers/pathways associated with a trait of interest.\n\nConclusionThe proposed DCD approach overcomes the limitations of previous statistical techniques and the issues associated with identifying differential sub-networks by use of community detection methods on the noisy DT graph. This is reflected in the superior performance of the DCD method with respect to various metrics like Precision, Accuracy, Kappa and Specificity. The code implementing proposed DCD method is available at https://sites.google.com/site/ raghvendramallmlresearcher/codes.

systems biology

A model of flux regulation in the cholesterol biosynthesis pathway: Immune mediated graduated flux reduction versus statin-like led stepped flux reduction

Graphical Abstract\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=37 SRC=\"FIGDIR/small/000380_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (14K):\norg.highwire.dtl.DTLVardef@197f26org.highwire.dtl.DTLVardef@1eac3f3org.highwire.dtl.DTLVardef@1e698f1org.highwire.dtl.DTLVardef@42fb4b_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIWe model the cholesterol biosynthesis pathway and its regulation\nC_LIO_LIThe innate immune response leads to a suppression of flux through the pathway\nC_LIO_LIStatin inhibitors show a different mode of suppression to the immune response\nC_LIO_LIStatin inhibitor suppression is less robust and less specific than immune suppression\nC_LI\n\nAsbtractThe cholesterol biosynthesis pathway has recently been shown to play an important role in the innate immune response to viral infection with host protection occurring through a coordinate down regulation of the enzymes catalyzing each metabolic step. In contrast, statin based drugs, which form the principle pharmaceutical agents for decreasing the activity of this pathway, target a single enzyme. Here, we build an ordinary differential equation model of the cholesterol biosynthesis pathway in order to investigate how the two regulatory strategies impact upon the behaviour of the pathway. We employ a modest set of assumptions: that the pathway operates away from saturation, that each metabolite is involved in multiple cellular interactions and that mRNA levels reflect enzyme concentrations. Using data taken from primary bone marrow derived macrophage cells infected with murine cytomegalovirus infection or treated with IFN{gamma}, we show that, under these assumptions, coordinate down regulation of enzyme activity imparts a graduated reduction in flux along the pathway. In contrast, modelling a statin-like treatment that achieves the same degree of down-regulation in cholesterol production, we show that this delivers a step change in flux along the pathway. The graduated reduction mediated by physiological coordinate regulation of multiple enzymes supports a mechanism that allows a greater level of specificity, altering cholesterol levels with less impact upon interactions branching from the pathway, than pharmacological step reductions. We argue that coordinate regulation is likely to show a long-term evolutionary advantage over single enzyme regulation. Finally, the results from our models have implications for future pharmaceutical therapies intended to target cholesterol production with greater specificity and fewer off target effects, suggesting that this can be achieved by mimicking the coordinated down-regulation observed in immunological responses.

Systems Biology

Quantifying the turnover of transcriptional subclasses of HIV-1-infected cells

HIV-1-infected cells in peripheral blood can be grouped into different transcriptional subclasses. Quantifying the turnover of these cellular subclasses can provide important insights into the viral life cycle and the generation and maintenance of latently infected cells. We used previously published data from five patients chronically infected with HIV-1 that initiated combination antiretroviral therapy (cART). Patient-matched PCR for unspliced and multiply spliced viral RNAs combined with limiting dilution analysis provided measurements of transcriptional profiles at the single cell level. Furthermore, measurement of intracellular transcripts and extracellular virion-enclosed HIV-1 RNA allowed us to distinguish productive from non-productive cells. We developed a mathematical model describing the dynamics of plasma virus and the transcriptional subclasses of HIV-1-infected cells. Fitting the model to the data allowed us to better understand the phenotype of different transcriptional subclasses and their contribution to the overall turnover of HIV-1 before and during cART. The average number of virus-producing cells in peripheral blood is small during chronic infection (25.7 cells ml-1). We find that 14.0%, 0.3% and 21.2% of infected cells become defectively, latently and persistently infected cells, respectively. Assuming that the infection is homogenous throughout the body, we estimate an average in vivo viral burst size of 2.1 x 104 virions per cell. Our study provides novel quantitative insights into the turnover and development of different subclasses of HIV-1-infected cells. The model predicts that the pool of latently infected cells becomes rapidly established during the first months of acute infection and continues to increase slowly during the first years of chronic infection. Having a detailed understanding of this process will be useful for the evaluation of viral eradication strategies that aim to deplete the latent reservoir of HIV-1.\n\nAuthor SummaryGaining a quantitative understanding of the development and turnover of different HIV-1-infected subpopulations of cells is crucial to improve the outcome of patients on combination antiretroviral therapy (cART). The population of latently infected cells is of particular interest as they represent the major barrier to a cure of HIV-1 infection. We developed a mathematical model that describes the dynamics of different transcriptionally active subclasses of HIV-1-infected cells and the viral load in peripheral blood. The model was fitted to previously published data from five chronically HIV-1-infected patients starting cART. This allowed us to estimate critical parameters of the within-host dynamics of HIV-1, such as the the number of virions produced by a single infected cell. The model further allowed investigation of HIV-1 dynamics during the acute phase. Computer simulations predict that latently infected cells become rapidly established during the first months of acute infection and continue to increase slowly during the first years of chronic infection. This illustrates the opportunity for strategies that aim to eradicate the virus during early cART as the pool of HIV-1 infected cells is substantially smaller during acute infection than during chronic infection.

Systems Biology

Population diversification in a yeast metabolic program promotes anticipation of environmental shifts

Delineating the strategies by which cells contend with combinatorial changing environments is crucial for understanding cellular regulatory organization. When presented with two carbon sources, microorganisms first consume the carbon substrate that supports the highest growth rate (e.g. glucose) and then switch to the secondary carbon source (e.g. galactose), a paradigm known as the Monod model. Sequential sugar utilization has been attributed to transcriptional repression of the secondary metabolic pathway, followed by activation of this pathway upon depletion of the preferred carbon source. In this work, we challenge this notion. Although Saccharomyces cerevisiae cells consume glucose before galactose, we demonstrate that the galactose regulatory pathway is activated in a fraction of the cell population hours before glucose is fully consumed. This early activation reduces the time required for the population to transition between the two metabolic programs and provides a fitness advantage that might be crucial in competitive environments. Importantly, these findings define a new paradigm for the response of microbial populations to combinatorial carbon sources.

Systems Biology

Measuring error rates in genomic perturbation screens: gold standards for human functional genomics

Technological advancement has opened the door to systematic genetics in mammalian cells. Genome-scale loss-of-function screens can assay fitness defects induced by partial gene knockdown, using RNA interference, or complete gene knockout, using new CRISPR techniques. These screens can reveal the basic blueprint required for cellular proliferation. Moreover, comparing healthy to cancerous tissue can uncover genes that are essential only in the tumor; these genes are targets for the development of specific anticancer therapies. Unfortunately, progress in this field has been hampered by offtarget effects of perturbation reagents and poorly quantified error rates in large-scale screens. To improve the quality of information derived from these screens, and to provide a framework for understanding the capabilities and limitations of CRISPR technology, we derive gold-standard reference sets of essential and nonessential genes, and provide a Bayesian classifier of gene essentiality that outperforms current methods on both RNAi and CRISPR screens. Our results indicate that CRISPR technology is more sensitive than RNAi, and that both techniques have nontrivial false discovery rates that can be mitigated by rigorous analytical methods.

Systems Biology

Towards a molecular systems model of coronary artery disease

Coronary artery disease (CAD) is a complex disease driven by myriad interactions of genetics and environmental factors. Traditionally, studies have analyzed only one disease factor at a time, providing useful but limited understanding of the underlying etiology. Recent advances in cost-effective and high-throughput technologies, such as single nucleotide polymorphism (SNP) genotyping, exome/genome sequencing, gene expression microarrays and metabolomics assays have enabled the collection of millions of data points in many thousands of individuals. In order to make sense of such omics data, effective analytical methods are needed. We review and highlight some of the main results in this area, focusing on integrative approaches that consider multiple modalities simultaneously. Such analyses have the potential to uncover the genetic basis of CAD, produce genomic risk scores (GRS) for disease prediction, disentangle the complex interactions underlying disease, and predict response to treatment.

Systems Biology

Exact Reconstruction of Gene Regulatory Networks using Compressive Sensing

BackgroundWe consider the problem of reconstructing a gene regulatory network structure from limited time series gene expression data, without any a priori knowledge of connectivity. We assume that the network is sparse, meaning the connectivity among genes is much less than full connectivity. We develop a method for network reconstruction based on compressive sensing, which takes advantage of the networks sparseness.\n\nResultsFor the case in which all genes are accessible for measurement, and there is no measurement noise, we show that our method can be used to exactly reconstruct the network. For the more general problem, in which hidden genes exist and all measurements are contaminated by noise, we show that our method leads to reliable reconstruction. In both cases, coherence of the model is used to assess the ability to reconstruct the network and to design new experiments. For each problem, a set of numerical examples is presented.\n\nConclusionsThe method provides a guarantee on how well the inferred graph structure represents the underlying system, reveals deficiencies in the data and model, and suggests experimental directions to remedy the deficiencies.

Systems Biology

A gene expression atlas of a bicoid-depleted Drosophila embryo reveals early canalization of cell fate

In developing embryos, gene regulatory networks canalize cells towards discrete terminal fates. We studied the behavior of the anterior-posterior segmentation network in Drosophila melanogaster embryos depleted of a key maternal input, bicoid (bcd), by building a cellular- resolution gene expression atlas containing measurements of 12 core patterning genes over 6 time points in early development. With this atlas, we determine the precise perturbation each cell experiences, relative to wild type, and observe how these cells assume cell fates in the perturbed embryo. The first zygotic layer of the network, consisting of the gap and terminal genes, is highly robust to perturbation: all combinations of transcription factor expression found in bcd depleted embryos were also found in wild type embryos, suggesting that no new cell fates were created even at this very early stage. All of the gap gene expression patterns in the trunk expand by different amounts, a feature that we were unable to explain using two simple models of the effect of bcd depletion. In the second layer of the network, depletion of bcd led to an excess of cells expressing both even skipped and fushi tarazu early in the blastoderm stage, but by gastrulation this overlap resolved into mutually exclusive stripes. Thus, following depletion of bcd, individual cells rapidly canalize towards normal cell fates in both layers of this gene regulatory network. Our gene expression atlas provides a high resolution picture of a classic perturbation and will enable further modeling of canalization in this transcriptional network.

Systems Biology