Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 469 records · Page 26Linked to original sources

eNetXplorer: an R package for the quantitative exploration of elastic net families for generalized linear models

SummarySystems biology analysis often involves building predictive models by selecting informative features from a large number of measurements. The elastic net for generalized linear models is a popular regression and feature selection method, particularly useful when the number of features is greater than the sample size or when there exist many correlated predictor variables. This package provides a quantitative, cross-validation based toolkit to evaluate elastic net models and to uncover correlates contributing to prediction. Feature importance is evaluated by flexible criteria using out-of-bag prediction performance assessed via user-defined quality functions. Statistical significance is assigned to each model by comparison to null models generated by permutations of sample labels; analogous approaches are used to assess significance for the contribution of individual features to prediction. This package fits linear, binomial (logistic) and multinomial models, and provides a set of standard plots, summary statistics and output tables. eNetXplorer enables quantitative, exploratory analysis to generate hypotheses on which features may be associated with biological phenotypes of interest, such as in the identification of biomarkers for therapeutic responsiveness.\n\nAvailability and implementationThe eNetXplorer R package is available under GPLv3 license at https://CRAN.R-project.org/package=eNetXplorer

systems biology

Saccharomyces cerevisiae displays a stable transcription start site landscape in multiple conditions

One of the fundamental processes that determine cellular fate is regulation of gene transcription. Understanding these regulatory processes is therefore essential for understanding cellular responses to changes in environmental conditions. At the core promoter, the regulatory region containing the transcription start site (TSS), all inputs regulating transcription are integrated. Here, we used Cap Analysis of Gene Expression (CAGE) to analyze the pattern of transcription start sites at four different environmental conditions (limited in ethanol, limited in nitrogen, limited in glucose and limited in glucose under anaerobic conditions) using the Saccharomyces cerevisiae strain CEN.PK113-7D. With this experimental setup we were able to show that the TSS landscape in yeast is stable at different metabolic states of the cell. We also show that the shape index, a characteristic feature of each TSS describing the spatial distribution of transcription initiation events, has a surprisingly strong negative correlation with the measured expression levels. Our analysis supplies a set of high quality TSS annotations useful for metabolic engineering and synthetic biology approaches in the industrially relevant laboratory strain CEN.PK113-7D, and provides novel insights into yeast TSS dynamics and gene regulation.

systems biology

Conceptual Confusion: the case of Epigenetics

The observations of phenotypic plasticity have stimulated the revival of epigenetics. Over the past 70 years the term has come in many colors and flavors, depending on the biological discipline and time period. The meanings span from Waddingtons \"epigenotype\" and \"epigenetic landscape\" to the molecular biologists \"epigenetic marks\" embodied by DNA methylation and histone modifications. Here we seek to quell the ambiguity of the name. First we place \"epigenetics\" in the various historical contexts. Then, by presenting the formal concepts of dynamical systems theory we show that the \"epigenetic landscape\" is more than a metaphor: it has specific mathematical foundations. The latter explains how gene regulatory networks produce multiple attractor states, the self-stabilizing patterns of gene activation across the genome that account for \"epigenetic memory\". This network dynamics approach replaces the reductionist correspondence of molecular epigenetic modifications with concept of the epigenetic landscape, by providing a concrete and crisp correspondence.

Systems Biology

Quantitative Characterization of Biological Age and Frailty Based on Locomotor Activity Records

We performed a systematic evaluation of the relationships between locomotor activity and signatures of frailty, morbidity, and mortality risks using physical activity records from the 2003 - 2006 National Health and Nutrition Examination Survey (NHANES) and UK BioBank (UKB). We proposed a statistical description of the locomotor activity tracks and transformed the provided time series into vectors representing physiological states for each participant. The Principal Components Analysis of the transformed data revealed a winding trajectory with distinct segments corresponding to subsequent human development stages. The extended linear phase starts from 35 40 years old and is associated with the exponential increase of mortality risks according to the Gompertz mortality law. We characterized the distance traveled along the aging trajectory as a natural measure of biological age and demonstrated its significant association with frailty and hazardous lifestyles, along with the remaining lifespan and healthspan of an individual. The biological age explained most of the variance of the log-hazard ratio that was obtained by fitting directly to mortality and the incidence of chronic diseases. Our findings highlight the intimate relationship between the supervised and unsupervised signatures of the biological age and frailty, a consequence of the low intrinsic dimensionality of the aging dynamics.

systems biology

Transcriptional burst initiation and polymerase pause release are key control points of transcriptional regulation

Transcriptional regulation occurs via changes to the rates of various biochemical processes. Sequencing-based approaches that average together many cells have suggested that polymerase binding and polymerase release from promoter-proximal pausing are two key regulated steps in the transcriptional process. However, single cell studies have revealed that transcription occurs in short, discontinuous bursts, suggesting that transcriptional burst initiation and termination might also be regulated steps. Here, we develop and apply a quantitative framework to connect changes in both Pol II ChIP-seq and single cell transcriptional measurements to changes in the rates of specific steps of transcription. Using a number of global and targeted transcriptional regulatory perturbations, we show that burst initiation rate is indeed a key regulated step, demonstrating that transcriptional activity can be frequency modulated. Polymerase pause release is a second key regulated step, but the rate of polymerase binding is not changed by any of the biological perturbations we examined. Our results establish an important role for transcriptional burst regulation in the control of gene expression.

systems biology

Large-scale analysis of the global transcriptional regulation of bacterial gene expression

Bacterial gene expression depends on the allocation of limited transcriptional resources provided a particular growth rate and growth condition. Early studies in a few genes suggested this global regulation to generate a unifying hyperbolic expression pattern. Here, we developed a large-scale method that generalizes these experiments to quantify the response to growth of over 700 genes that a priori do not exhibit any specific control. We distinguish a core subset following a promoter-specific hyperbolic response. Within this group, we sort genes with regard to their responsiveness to the global regulatory program to show that those with a particularly sensitive linear response are located near the origin of replication. We then find evidence that this genomic architecture is biologically significant by examining position conservation of E. coli genes in 100 bacteria. The response to the transcriptional resources of the cell results consequently in an additional feature contributing to bacterial genome organization.

systems biology

Optimal homotopy analysis of a chaotic HIV-1 model incorporating AIDS-related cancer cells

The studies of nonlinear models in epidemiology have generated a deep interest in gaining insight into the mechanisms that underlie AIDS-related cancers, providing us with a better understanding of cancer immunity and viral oncogenesis. In this article, we analyse an HIV-1 model incorporating the relations between three dynamical variables: cancer cells, healthy CD4+ T lymphocytes and infected CD4+ T lymphocytes. Recent theoretical investigations indicate that these cells interactions lead to different dynamical outcomes, for instance to periodic or chaotic behavior. Firstly, we analytically prove the boundedness of the trajectories in the systems attractor. The complexity of the coupling between the dynamical variables is quantified using observability indices. Our calculations reveal that the highest observable variable is the population of cancer cells, thus indicating that these cells could be monitored in future experiments in order to obtain time series for attractors reconstruction. We identify different dynamical behaviors of the system varying two biologically meaningful parameters: r1, representing the uncontrolled proliferation rate of cancer cells, and k1, denoting the immune systems killing rate of cancer cells. The maximum Lyapunov exponent is computed to identify the chaotic regimes. Considering very recent developments in the literature related to the homotopy analysis method (HAM), we construct the explicit series solution of the cancer model and focus our analysis on the dynamical variable with the highest observability index. An optimal homotopy analysis approach is used to improve the computational efficiency of HAM by means of appropriate values for the convergence control parameter, which greatly accelerate the convergence of the series solution.

systems biology

Subcellular Time Series Modeling of Heterogeneous Cell Protrusion

In this paper, a new biological modeling approach is proposed for predicting complex heterogeneous subcellular behaviors. Cell protrusion which initiates cell migration has a significant amount of subcellular heterogeneity in micrometer length and minute time scales. It is driven by actin polymerization, e.g., pushing the plasma membrane forward, and then regulated by a multitude of actin regulators. While mathematical modeling is central to system-level understandings of cell protrusion, most of the modeling is based on the ensemble average of actin regulator dynamics at the cellular or population levels, preventing from capturing the heterogeneous cellular activities. With these in mind, a systematic modeling framework is proposed in this paper for predicting velocities of heterogeneous protrusion of migrating cells driven by multiple molecular mechanisms. The modeling framework is developed through the integration of the multiple AutoRegressive eXogenous (ARX) models employing probability density input variables. Unlike conventional ARX models, it provides an effective framework for modeling heterogeneous subcellular behaviors with complex nonlinearities and uncertainties of dynamic systems. To train and validate the proposed model, numerous subcellular time series are extracted from time-lapse movies of migrating PtK1 cells using spinning disk confocal microscope: The current edge velocities and fluorescent intensities of mDia1, actin at the leading edge are used as the input while the future cell edge velocities are selected as an output. It is demonstrated that the proposed approach is highly effective in predicting the future trends of heterogeneous cell protrusion. In particular, by capturing the various multiple activities from the dataset, it is expected that it would improve the understanding of the molecular mechanism underlying cellular and subcellular heterogeneity.

systems biology

ModuleDiscoverer: Identification of regulatory modules in protein-protein interaction networks.

The identification of disease associated modules based on protein-protein interaction networks (PPINs) and gene expression data has provided new insights into the mechanistic nature of diverse diseases. A major problem hampering their identification is the detection of protein communities within large-scale, whole-genome PPINs. Current strategies solve the maximal clique enumeration (MCE) problem, i.e., the enumeration of all non-extendable groups of proteins, where each pair of proteins is connected by an edge. The MCE problem however is non-deterministic polynomial time hard and can thus be computationally overwhelming for large-scale, whole-genome PPINs.\n\nWe present ModuleDiscoverer, a novel approach for the identification of regulatory modules from PPINs in conjunction with gene-expression data. ModuleDiscoverer is a heuristic that approximates the community structure underlying PPINs. Based on a high-confidence PPIN of Rattus norvegicus and publicly available gene expression data we apply our algorithm to identify the regulatory module of a rat-model of diet induced non-alcoholic steatohepatitis (NASH). We validate the module using single-nucleotide polymorphism data from independent genome-wide association studies. Structural analysis of the module reveals 10 sub-modules. These sub-modules are associated with distinct biological functions and pathways that are relevant to the pathological and clinical situation in NASH.\n\nModuleDiscoverer is freely available upon request from the corresponding author.

systems biology

CNNC: Convolutional Neural Networks for Co-Expression Analysis

Several methods were developed to mine gene-gene relationships from expression data. Examples include correlation and mutual information methods for co-expression analysis, clustering and undirected graphical models for functional assignments and directed graphical models for pathway reconstruction. Using a novel encoding for gene expression data, followed by deep neural networks analysis, we present a framework that can successfully address all these diverse tasks. We show that our method, CNNC, improves upon prior methods in tasks ranging from predicting transcription factor targets to identifying disease related genes to causality inference. CNNCs encoding provides insights about some of the decisions it makes and their biological basis. CNNC is flexible and can easily be extended to integrate additional types of genomics data leading to further improvements in its performance.\n\nSupporting website with software and data: https://github.com/xiaoyeye/CNNC.

systems biology

Memory sequencing reveals heritable single cell gene expression programs associated with distinct cellular behaviors

Non-genetic factors can cause individual cells to fluctuate substantially in gene expression levels over time. Yet it remains unclear whether these fluctuations can persist for much longer than the time of one cell division. Current methods for measuring gene expression in single cells mostly rely on single time point measurements, making the duration of gene expression fluctuations or cellular memory difficult to measure. Here, we report a method combining Luria and Delbrucks fluctuation analysis with population-based RNA sequencing (MemorySeq) for identifying genes transcriptome-wide whose fluctuations persist for several cell divisions. MemorySeq revealed multiple gene modules that are expressed together in rare cells within otherwise homogeneous clonal populations. Further, we found that these rare cell subpopulations are associated with biologically distinct behaviors, such as the ability to proliferate in the face of anti-cancer therapeutics, in different cancer cell lines. The identification of non-genetic, multigenerational fluctuations has the potential to reveal new forms of biological memory at the level of single cells and suggests that non-genetic heritability of cellular state may be a quantitative property.

systems biology

Network-based prediction of protein interactions

As biological function emerges through interactions between a cells molecular constituents, understanding cellular mechanisms requires us to catalogue all physical interactions between proteins [1-4]. Despite spectacular advances in high-throughput mapping, the number of missing human protein-protein interactions (PPIs) continues to exceed the experimentally documented interactions [5, 6]. Computational tools that exploit structural, sequence or network topology information are increasingly used to fill in the gap, using the patterns of the already known interactome to predict undetected, yet biologically relevant interactions [7-9]. Such network-based link prediction tools rely on the Triadic Closure Principle (TCP) [10-12], stating that two proteins likely interact if they share multiple interaction partners. TCP is rooted in social network analysis, namely the observation that the more common friends two individuals have, the more likely that they know each other [13, 14]. Here, we offer direct empirical evidence across multiple datasets and organisms that, despite its dominant use in biological link prediction, TCP is not valid for most protein pairs. We show that this failure is fundamental - TCP violates both structural constraints and evolutionary processes. This understanding allows us to propose a link prediction principle, consistent with both structural and evo-lutionary arguments, that predicts yet uncovered protein interactions based on paths of length three (L3). A systematic computational cross-validation shows that the L3 principle significantly outperforms existing link prediction methods. To experimentally test the L3 predictions, we perform both large-scale high-throughput and pairwise tests, finding that the predicted links test positively at the same rate as previously known interactions, suggesting that most (if not all) predicted interactions are real. Combining L3 predictions with experimen-tal tests provided new interaction partners of FAM161A, a protein linked to retinitis pigmentosa, offering novel insights into the molecular mechanisms that lead to the disease. Because L3 is rooted in a fundamental biological principle, we expect it to have a broad applicability, enabling us to better understand the emergence of biological function under both healthy and pathological conditions.\n\nSummaryWe unveil a fundamental organizing principle of biological networks and demonstrate its predictive power for uncovering novel protein interactions.

systems biology

Intercellular signaling network underlies biological time across multiple temporal scales

MotivationCellular, physiological and molecular processes must be organized and regulated across multiple time domains throughout the lifespan of an organism. The technological revolution in molecular biology has led to the identification of numerous genes implicated in the regulation of diverse temporal biological processes. However, it is natural to question whether there is an underlying regulatory network governing multiple timescales simultaneously.\n\nResultsUsing queries of relevant databases and literature searches, a single dense multiscale temporal regulatory network was identified involving core sets of genes that regulate circadian, cell cycle, and aging processes. The network was highly enriched for genes involved in signal transduction (P = 1.82e-82), with p53 and its regulators such as p300 and CREB binding protein forming key hubs, but also for genes involved in metabolism (P = 6.07e-127) and cellular response to stress (P = 1.56e-93). These results suggest an intertwined molecular signaling network that affects biological time across multiple temporal scales in response to environmental stimuli and available resources.\n\nContactjoshua.millstein@usc.edu\n\nSupplementary informationSupplementary data are available online.

systems biology

Flux-dependent graphs for metabolic networks

Cells adapt their metabolic fluxes in response to changes in the environment. We present a frame-work for the systematic construction of flux-based graphs derived from organism-wide metabolic networks. Our graphs encode the directionality of metabolic fluxes via edges that represent the flow of metabolites from source to target reactions. The methodology can be applied in the absence of a specific biological context by modelling fluxes probabilistically, or can be tailored to different environ-mental conditions by incorporating flux distributions computed through constraint-based approaches such as Flux Balance Analysis. We illustrate our approach on the central carbon metabolism of Escherichia coli and on a metabolic model of human hepatocytes. The flux-dependent graphs under various environmental conditions and genetic perturbations exhibit systemic changes in their topo-logical and community structure, which capture the re-routing of metabolic fluxes and the varying importance of specific reactions and pathways. By integrating constraint-based models and tools from network science, our framework allows the study of context-specific metabolic responses at a system level beyond standard pathway descriptions.

systems biology

Increasing consensus of context-specific metabolic models by integrating data-inferred cell functions

Genome-scale metabolic models provide a valuable context for analyzing data from diverse high-throughput experimental techniques. Models can quantify the activities of diverse pathways and cellular functions. Since some metabolic reactions are only catalyzed in specific environments, several algorithms exist that build context-specific models. However, these methods make differing assumptions that influence the content and associated predictive capacity of resulting models, such that model content varies more due to methods used than cell types. Here we overcome this problem with a novel framework for inferring the metabolic functions of a cell before model construction. For this, we curated a list of metabolic tasks and developed a framework to infer the activity of these functionalities from transcriptomic data. We protected the data-inferred tasks during the implementation of diverse context-specific model extraction algorithms for 44 cancer cell lines. We show that the protection of data-inferred metabolic tasks decreases the variability of models across extraction methods. Furthermore, resulting models better capture the actual biological variability across cell lines. This study highlights the potential of using biological knowledge, inferred from omics data, to obtain a better consensus between existing extraction algorithms. It further provides guidelines for the development of the next-generation of data contextualization methods.

systems biology

Quasi-universality in single-cell sequencing data

The development of single-cell technologies provides the opportunity to identify new cellular states and reconstruct novel cell-to-cell relationships. Applications range from understanding the transcriptional and epigenetic processes involved in metazoan development to characterizing distinct cells types in heterogeneous populations like cancers or immune cells. However, analysis of the data is impeded by its unknown intrinsic biological and technical variability together with its sparseness; these factors complicate the identification of true biological signals amidst artifact and noise. Here we show that, across technologies, roughly 95% of the eigenvalues derived from each single-cell data set can be described by universal distributions predicted by Random Matrix Theory. Interestingly, 5% of the spectrum shows deviations from these distributions and present a phenomenon known as eigenvector localization, where information tightly concentrates in groups of cells. Some of the localized eigenvectors reflect underlying biological signal, and some are simply a consequence of the sparsity of single cell data; roughly 3% is artifactual. Based on the universal distributions and a technique for detecting sparsity induced localization, we present a strategy to identify the residual 2% of directions that encode biological information and thereby denoise single-cell data. We demonstrate the effectiveness of this approach by comparing with standard single-cell data analysis techniques in a variety of examples with marked cell populations.

systems biology

Investigating the relation between stochastic differentiation and homeostasis in intestinal crypts via multiscale modeling

Colorectal tumors originate and develop within intestinal crypts. Even though some of the essential phenomena that characterize crypt structure and dynamics have been effectively described in the past, the relation between the differentiation process and the overall crypt homeostasis is still partially understood. We here investigate this relation and other important biological phenomena by introducing a novel multiscale model that combines a morphological description of the crypt with a gene regulation model: the emergent dynamical behavior of the underlying gene regulatory network drives cell growth and differentiation processes, linking the two distinct spatio-temporal levels. The model relies on a few a priori assumptions, yet accounting for several key processes related to crypt functioning, such as: dynamic gene activation patterns, stochastic differentiation, signaling pathways ruling cell adhesion properties, cell displacement, cell growth, mitosis, apoptosis and the presence of biological noise.\n\nWe show that this modeling approach captures the major dynamical phenomena that characterize the regular physiology of crypts, such as cell sorting, coordinate migration, dynamic turnover, stem cell niche maintenance and clonal expansion. All in all, the model suggests that the process of stochastic differentiation might be sufficient to drive the crypt to homeostasis, under certain crypt configurations. Besides, our approach allows to make precise quantitative inferences that, when possible, were matched to the current biological knowledge and it permits to investigate the role of gene-level perturbations, with reference to cancer development. We also remark the theoretical framework is general and may applied to different tissues, organs or organisms.

Systems Biology

Information Processing by Simple Molecular Motifs and Susceptibility to Noise

Biological organisms rely on their ability to sense and respond appropriately to their environment. The molecular mechanisms that facilitate these essential processes are however subject to a range of random effects and stochastic processes, which jointly affect the reliability of information transmission between receptors and e.g. the physiological downstream response. Information is mathematically defined in terms of the entropy; and the extent of information flowing across an information channel or signalling system is typically measured by the \"mutual information\", or the reduction in the uncertainty about the output once the input signal is known. Here we quantify how extrinsic and intrinsic noise affect the transmission of simple signals along simple motifs of molecular interaction networks. Even for very simple systems the effects of the different sources of variability alone and in combination can give rise to bewildering complexity. In particular extrinsic variability is apt to generate \"apparent\" information that can in extreme cases mask the actual information that for a single system would flow between the different molecular components making up cellular signalling pathways. We show how this artificial inflation in apparent information arises and how the effects of different types of noise alone and in combination can be understood.

Systems Biology