Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Mechanistic modelling and Bayesian inference elucidates the variable dynamics of double-strand break repair

DNA double-strand breaks are lesions that form during metabolism, DNA replication and exposure to mutagens. When a double-strand break occurs one of a number of repair mechanisms is recruited, all of which have differing propensities for mutational events. Despite DNA repair being of crucial importance, the relative contribution of these mechanisms and their regulatory interactions remain to be fully elucidated. Understanding these mutational processes will have a profound impact on our knowledge of genomic instability, with implications across health, disease and evolution. Here we present a new method to model the combined activation of non-homologous end joining, single strand annealing and alternative end joining, following exposure to ionizing radiation. We use Bayesian statistics to integrate eight biological data sets of double-strand break repair curves under varying genetic knockouts and confirm that our model is predictive by re-simulating and comparing to additional data. Analysis of the model suggests that there are at least three disjoint modes of repair, which we assign as fast, slow and intermediate. Our results show that when multiple data sets are combined, the rate for intermediate repair is variable amongst genetic knockouts. Further analysis suggests that the ratio between slow and intermediate repair depends on the presence or absence of DNA-PKcs and Ku70, which implies that non-homologous end joining and alternative end joining are not independent. Finally, we consider the proportion of double-strand breaks within each mechanism as a time series and predict activity as a function of repair rate. We outline how our insights can be directly tested using imaging and sequencing techniques and conclude that there is evidence of variable dynamics in alternative repair pathways. Our approach is an important step towards providing a unifying theoretical framework for the dynamics of DNA repair processes.

Systems Biology

Transcriptomes and Raman spectra are linked linearly through a shared low-dimensional subspace

Raman spectroscopy is an imaging technique that can reflect whole-cell molecular compositions in vivo, and has been applied recently in cell biology to characterize different cell types and states. However, due to the complex molecular compositions and spectral overlaps, the interpretation of cellular Raman spectra have remained unclear. In this report, we compared cellular Raman spectra to transcriptomes of Schizosaccharomyces pombe and Escherichia coli, and provide firm evidence that they can be computationally connected and interpreted. Specifically, we find that the dimensions of high-dimensional Raman spectra and transcriptomes measured by RNA-seq can be effectively reduced and connected linearly through a shared low-dimensional subspace. Accordingly, we were able to reconstruct global gene expression profiles by applying the calculated transformation matrix to Raman spectra, and vice versa. Strikingly, highly expressed ncRNAs contributed to the Raman-transcriptome linear correspondence more significantly than mRNAs in S. pombe, which implies their major role in coordinating molecular compositions. This compatibility between whole-cell Raman spectra and transcriptomes marks an important and promising step towards establishing spectroscopic live-cell omics studies.

systems biology

BNMCMC: a software for learning and visualizing Bayesian networks using MCMC methods

MotivationBayesian networks (BNs) are widely used to model biological networks from experimental data. Many software packages exist to infer BN structures, but the chance of getting trapped in local optima is a common challenge. Some recently developed Markov Chain Monte Carlo (MCMC) samplers called the Neighborhood sampler (NS) and Hit-and-Run (HAR) sampler, have shown great potential to substantially avoid this problem compared to the standard Metropolis-Hastings (MH) sampler.\n\nResultsWe have developed a software called BNMCMC for inferring and visualizing BNs from given datasets. This software runs NS, HAR and MH samplers using a discrete Bayesian model. The main advantage of BNMCMC is that it exploits adaptive techniques to efficiently explore BN space and evaluate the posterior probability of candidate BNs to facilitate large-scale network inference.\n\nAvailabilityBNMCMC is implemented with C#.NET, ASP.NET, Jquery, Javascript and D3.js. The standalone version (BN visualization missing) available for downloading at https://sourceforge.net/projects/bnmcmc/, where the user-guide and an example file are provided for a simulation. A dedicated BNMCMC web server will be launched soon feature a physics-based BN visualization technique.\n\nContactakm.azad@unsw.edu.au

systems biology

Identification of microRNA clusters cooperatively acting on Epithelial to Mesenchymal Transition in Triple Negative Breast Cancer

MicroRNAs play important roles in many biological processes. Their aberrant expression can have oncogenic or tumor suppressor function directly participating to carcinogenesis, malignant transformation, invasiveness and metastasis. Indeed, miRNA profiles can distinguish not only between normal and cancerous tissue but they can also successfully classify different subtypes of a particular cancer. Here, we focus on a particular class of transcripts encoding polycistronic miRNA genes that yields multiple miRNA components. We describe clustered MiRNA Master Regulator Analysis (ClustMMRA), a fully redesigned release of the MMRA computational pipeline (MiRNA Master Regulator Analysis), developed to search for clustered miRNAs potentially driving cancer molecular subtyping. Genomically clustered miRNAs are frequently co-expressed to target different components of pro-tumorigenic signalling pathways. By applying ClustMMRA to breast cancer patient data, we identified key miRNA clusters driving the phenotype of different tumor subgroups. The pipeline was applied to two independent breast cancer datasets, providing statistically concordant results between the two analysis. We validated in cell lines the miR-199/miR-214 as a novel cluster of miRNAs promoting the triple negative subtype phenotype through its control of proliferation and EMT.

systems biology

Noise Analysis in Biochemical Complex Formation

Several biological functions are carried out via complexes that are formed via multimerization of either a single species (homomers) or multiple species (heteromers). Given functional relevance of these complexes, it is arguably desired to maintain their level at a set point and minimize fluctuations around it. Here we consider two simple models of complex formation - one for homomer and another for heteromer of two species - and analyze how important model parameters affect the noise in complex level. In particular, we study effects of (i) sensitivity of the complex formation rate with respect to constituting species abundance, and (ii) relative stability of the complex as compared with that of the constituents. By employing an approximate moment analysis, we find that for a given steady state level, there is an optimal sensitivity that minimizes noise (quantified by fano-factor; variance/mean) in the complex level. Furthermore, the noise becomes smaller if the complex is less stable than its constituents. Finally, for the heteromer case, our findings show that noise is enhanced if the complex is comparatively more sensitive to one constituent. We briefly discuss implications of our result for general complex formation processes.

systems biology

Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer

BackgroundDe novo inference of clinically relevant gene function relationships from tumor RNA-seq remains a challenging task. Current methods typically either partition patient samples into a few subtypes or rely upon analysis of pairwise gene correlations (co-expression) that will miss some groups in noisy data. Leveraging higher dimensional information can be expected to increase the power to discern targetable pathways, but this is commonly thought to be an intractable computational problem.\n\nMethodsIn this work we adapt a recently developed machine learning algorithm, CorEx, that efficiently optimizes over multivariate mutual information for sensitive detection of complex gene relationships. The algorithm can be iteratively applied to generate a hierarchy of latent factors. Patients are stratified relative to each factor and combinatoric survival analyses are performed and interpreted in the context of biological function annotations and protein network interactions that might be utilized to match patients to multiple therapies.\n\nResultsAnalysis of ovarian tumor RNA-seq samples demonstrates the algorithms power to infer well over one hundred biologically interpretable gene cohorts, several times more than standard methods such as hierarchical clustering and k-means. The CorEx factor hierarchy is also informative, with related but distinct gene clusters grouped by upper nodes. Some latent factors correlate with patient survival, including one for a pathway connected with the epithelial-mesenchymal transition in breast cancer that is regulated by a potentially druggable microRNA. Further, combinations of factors lead to a synergistic survival advantage in some cases.\n\nConclusionsIn contrast to studies that attempt to partition patients into a small number of subtypes (typically 4 or fewer) for treatment purposes, our approach utilizes subgroup information for combinatoric transcriptional phenotyping. Considering only the 66 gene expression groups that are both found to have significant Gene Ontology enrichment and are small enough to indicate specific drug targets implies a computational phenotype for ovarian cancer that allows for 366 possible patient profiles, enabling truly personalized treatment. The findings here demonstrate a new technique that sheds light on the complexity of gene expression dependencies in tumors and could eventually enable the use of patient RNA-seq profiles for selection of personalized and effective cancer treatments.

Systems Biology

Two stochastic processes shape diverse senescence patterns in a single-cell organism

Despite advances in aging research, a multitude of aging models, and empirical evidence for diverse senescence patterns, understanding is lacking of the biological processes that shape senescence, both for simple and complex organisms. We show that for a isogenic Escherichia coli bacterial population senescence results from two stochastic processes. A primary random deterioration process within the cell, such as generated by random accumulation of damage, leads to an exponential increase in mortality early in life followed by a late age mortality plateau; a secondary process of stochastic asymmetric transmission of an unknown factor at cell fission influences mortality. This second process is required to explain the difference between the classical mortality plateaus detected for young mothers offspring and the near non-senescence of old mothers offspring as well as the lack of a mother offspring correlation in age at death. We observed that life span is predominantly determined by underlying stochastic stage dynamics. Our findings support models based on stage-specific actions of alleles for the evolution of senescence. This support might be surprising since these models that have not specifically been developed in the context of simple, single cell organisms. We call for exploration of similar stochastic influences beyond simple organisms.

systems biology

Comparison of single gene and module-based methods for modeling gene regulatory networks

Gene regulatory networks describe the regulatory relationships among genes, and developing methods for reverse engineering these networks are an ongoing challenge in computational biology. The majority of the initially proposed methods for gene regulatory network discovery create a network of genes and then mine it in order to uncover previously unknown regulatory processes. More recent approaches have focused on inferring modules of co-regulated genes, linking these modules with regulator genes and then mining them to discover new molecular biology.\n\nIn this work we analyze module-based network approaches to build gene regulatory networks, and compare their performance to the well-established single gene network approaches. In particular, we focus on the problem of linking genes with known regulatory genes. First, modules are created iteratively using a regression approach that links co-expressed genes with few regulatory genes. After the modules are built, we create bipartite graphs to identify a set of target genes for each regulatory gene. We analyze several methods for uncovering these modules and show that a variational Bayes approach achieves significant improvement with respect to previously used methods for module creation on both simulated and real data. We also perform a topological and gene set enrichment analysis and compare several module-based approaches to single gene network approaches where a graph is built from the gene expression profiles without clustering genes in modules. We show that the module-based approach with variational Bayes outperforms all other methods and creates regulatory networks with a significantly higher rate of enriched molecular pathways.\n\nThe code is written in R and can be downloaded from https://github.com/mikelhernaez/linker.

systems biology

An analytical approach to bistable biological circuit discrimination using real algebraic geometry

Biomolecular circuits with two distinct and stable steady states have been identified as essential components in a wide range of biological networks, with a variety of mechanisms and topologies giving rise to their important bistable property. Understanding the differences between circuit implementations is an important question, particularly for the synthetic biologist faced with determining which bistable circuit design out of many is best for their specific application. In this work we explore the applicability of Sturms theorem--a tool from 19th-century real algebraic geometry--to comparing \"functionally equivalent\" bistable circuits without the need for numerical simulation. We first consider two genetic toggle variants and two different positive feedback circuits, and show how specific topological properties present in each type of circuit can serve to increase the size of the regions of parameter space in which they function as switches. We then demonstrate that a single competitive monomeric activator added to a purely-monomeric (and otherwise monostable) mutual repressor circuit is sufficient for bistability. Finally, we compare our approach with the Routh-Hurwitz method and derive consistent, yet more powerful, parametric conditions. The predictive power and ease of use of Sturms theorem demonstrated in this work suggests that algebraic geometric techniques may be underutilized in biomolecular circuit analysis.

Systems Biology

GOcats: A tool for categorizing Gene Ontology into subgraphs of user-defined concepts

Gene Ontology is used extensively in scientific knowledgebases and repositories to organize the wealth of available biological information. However, interpreting annotations derived from differential gene lists is difficult without manually sorting into higher-order categories. To address these issues, we present GOcats, a novel tool that organizes the Gene Ontology (GO) into subgraphs representing user-defined concepts, while ensuring that all appropriate relations are congruent with respect to scoping semantics. We tested GOcats performance using subcellular location categories to mine annotations from GO-utilizing knowledgebases and evaluating their accuracy against immunohistochemistry datasets in the Human Protein Atlas (HPA).\n\nIn comparison to mappings generated from UniProts controlled vocabulary and from GO slims via OWLTools Map2Slim, GOcats outperforms these methods without reliance on a human-curated set of GO terms. By identifying and properly defining relations with respect to semantic scope, GOcats can use traditionally problematic relations without encountering erroneous term mapping. We applied GOcats in the comparison of HPA-sourced knowledgebase annotations to experimentally-derived annotations provided by HPA directly. During the comparison, GOcats improved correspondence between the annotation sources by adjusting semantic granularity. Utilized in this way, GOcats can perform an accurate knowledgebase-level evaluation of curated HPA-based annotations.

systems biology

Towards the quantitative characterization of piglets’ robustness to weaning: A modelling approach

Weaning is a critical transition phase in swine production in which piglets must cope with different stressors that may affect their health. During this period, the prophylactic use of antibiotics is still frequent to limit piglet morbidity, which raises both economic and public health concerns such as the appearance of antimicrobial-resistant microbes. With the interest of developing tools for assisting health and management decisions around weaning, it is key to provide robustness indexes that inform on the animals capacity to endure the challenges associated to weaning. This work aimed at developing a modelling approach for facilitating the quantification of piglet resilience to weaning. We monitored 325 Large White pigs weaned at 28 days of age and further housed and fed conventionally during the post-weaning period without antibiotic administration. Body weight and diarrhoea scores were recorded before and after weaning, and blood was sampled at weaning and one week later for collecting haematological data. We constructed a dynamic model based on the Gompertz-Makeham law to describe live weight trajectories during the first 75 days after weaning following the rationale that the animal response is partitioned in two time windows (a perturbation and a recovery window). Model calibration was performed for each animal. Our results show that the transition time between the two time windows, as well as the weight trajectories are characteristic for each individual. The model captured the weight dynamics of animals at different degrees of perturbation, with an average coefficient of determination of 0.99, and a concordance correlation coefficient of 0.99. The utility of the model is that it provides biological parameters that inform on the amplitude and length of perturbation, and the rate of animal recovery. Our rationale is that the dynamics of weight inform on the capability of the animal to cope with the weaning disturbance. Indeed, there were significant correlations between model parameters and individual diarrhoea scores and haematological traits. Overall, the parameters of our model can be useful for constructing weaning robustness indexes by using exclusively the growth curves. We foresee that this modelling approach will provide a step forward in the quantitative characterization of robustness.\n\nImplicationsThe quantitative characterization of animal robustness at weaning is a key step for management strategies to improve health and welfare. This characterization is also instrumental for the further design of selection strategies for productivity and robustness. Within a precision livestock farming optic, this study develops a mathematical modelling approach to describe the body weight of piglets from weaning with the rationale that weight trajectories provide central information to quantify the capability of the animal to cope with the weaning disturbance.

systems biology

Algorithmic biosynthesis of eukaryotic glycans

An algorithm converts inputs to corresponding unique outputs through a sequence of actions. Algorithms are used as metaphors for complex biological processes such as organismal development. Here we make this metaphor rigorous for glycan biosynthesis. Glycans are branched sugar oligomers that are attached to cell-surface proteins and convey cellular identity. Eukaryotic O-glycans are synthesized by collections of enzymes in Golgi compartments. A compartment can stochastically convert a single input oligomer to a heterogeneous set of possible output oligomers; yet a given type of protein is invariably associated with a narrow and reproducible glycan oligomer profile. Here we resolve this paradox by borrowing from the theory of algorithmic self-assembly. We rigorously enumerate the sources of glycan microheterogeneity: incomplete oligomers via early exit from the reaction compartment; tandem repeat oligomers via runaway reactions; and competing oligomer fates via divergent reactions. We demonstrate how to diagnose and eliminate each of these, thereby obtaining \"algorithmic compartments\" that convert inputs to corresponding unique outputs. Given an input and a target output we either prove that the output cannot be algorithmically synthesized from the input, or explicitly construct an ordered series of algorithmic compartments that achieves this synthesis. Our theoretical analysis allows us to infer the causes of non-algorithmic microheterogeneity and species-specific diversity in real glycan datasets.

systems biology

Analysis of the transcriptional logic governing differential spatial expression in Hh target genes

This work provides theoretical tools to analyse the transcriptional effects of certain biochemical mechanisms (i.e. affinity and cooperativity) that have been proposed in previous literature to explain the differential spatial expression of Hedgehog target genes involved in Drosophila development. Specifically we have focused on the expression of decapentaplegic and patched. The transcription of these genes is believed to be controlled by opposing gradients of the activator and repressor forms of the transcription factor Cubitus interruptus (Ci). This study is based on a thermodynamic approach, which provides expression rates for these genes. These expression rates are controlled by transcription factors which are competing and cooperating for common binding sites. We have made mathematical representations of the different expression rates which depend on multiple factors and variables. The expressions obtained with the model have been refined to produce simpler equivalent formulae which allow for their mathematical analysis. Thanks to this, we can evaluate the correlation between the different interactions involved in transcription and the biological features observed at tissular level. These mathematical models can be applied to other morphogenes to help understand the complex transcriptional logic of opposing activator and repressor gradients.\n\nAuthor summaryMorphogenic differentiation is a complex process that involves emission, reception and cellular response to different signals. It is well known that the same morphogenic signal can give rise to different cellular transcriptional responses that usually depend, among other factors, on transcription factors. In concordance with the activator threshold model, classically it has been distinguished between high and low threshold target genes in order to explain how cells receiving the same signal can activate different genes. However, in particular cases where the transcription is controlled by two opposing transcription factors, it has been tested that this logic is not valid. This motivates the necessity for describing new theoretical models in order to understand better these cellular responses. By a theoretical analysis we have deduced different versions of transcriptional logic that are significantly determined by how the opposing transcription factors cooperate between them in the transcription process. We have also tested these different scenarios focussing on the Drosophila Hh target genes, and we have reproduced similar conclusions to the ones obtained by other methodologies.

systems biology

Towards a unified resource for transcriptional regulation in Escherichia coli K-12: Incorporating high throughput-generated binding data within the classic framework of regulation of initiation of transcription in RegulonDB.

Our understanding of the regulation of gene expression has been strongly benefited by the availability of high throughput technologies that enable questioning the whole genome for the binding of specific transcription factors and expression profiles. In the case of genome models, such as Escherichia coli K-12, this knowledge needs to be integrated with the legacy of accumulated genetics and molecular biology pre-genomic knowledge in order to attain deeper levels in the understanding of their biology. In spite of the several repositories and curated databases, there is no effort, nor electronic site yet, to comprehensively integrate the available knowledge from all these different sources around the regulation of gene expression of E. coli K-12. In this paper, we describe a first effort to expand RegulonDB, the database containing the rich legacy of decades of classic molecular biology experiments supporting what we know about gene regulation and operon organization in E. coli K-12, to include the genome-wide data set collections from 25 ChIP and 18 gSELEX publications, respectively, in addition to around 60 expression profiles used in their curation. Three essential features for the integration of this information coming from different methodological approaches are; first, a controlled vocabulary within an ontology for precisely defining growth conditions, second, the criteria to separate elements with enough evidence to consider them involved in gene regulation from isolated sites, and third, an expanded computational model supporting this knowledge. Altogether, this constitutes the basis for adequately gathering and enabling the comparisons and integration strongly needed to manage and access such wealth of knowledge. This version of RegulonBD is a first step toward what should become the unifying access point for current and future knowledge on gene regulation in E. coli K-12. Furthermore, this model platform and associated methodologies and criteria, can well be emulated for gathering knowledge on other microbial organisms.

systems biology

DensityPath: a level-set algorithm to visualize and reconstruct cell developmental trajectories for large-scale single-cell RNAseq data

Cell fates are determined by transition-states which occur during complex biological pro-cesses such as proliferation and differentiation. The advance in single-cell RNA sequencing (scRNAseq) provides the snapshots of single cell transcriptomes, thus offering an essential opportunity to study such complex biological processes. Here, we introduce a novel algorithm, DensityPath, which visualizes and reconstructs the underlying cell developmental trajectories for large-scale scRNAseq data. DensityPath has three merits. Firstly, by adopting the nonlinear dimension reduction algorithm elastic embedding, DensityPath reveals the intrinsic structures of the data. Secondly, by applying the powerful level set clustering method, DensityPath extracts the separate high density clusters of representative cell states (RCSs) from the single cell multimodal density landscape of gene expression space, enabling it to handle the heterogeneous scRNAseq data elegantly and accurately. Thirdly, DensityPath constructs cell state-transition path by finding the geodesic minimum spanning tree of the RCSs on the surface of the density landscape, making it more computationally efficient and accurate for large-scale dataset. The cell state-transition path constructed by DensityPath has the physical interpretation as the minimum-transition-energy (least-cost) path. We demonstrate that DensityPath is capable of identifying complex cell development trajectories with bifurcating and trifurcating branches on the human preimplantation embryos. We demonstrate that DensityPath is robust and has high accuracy of pseudotime calculation and branch assignment on the real scRNAseq as well as simulated datasets.

systems biology

IndeCut evaluates performance of network motif discovery algorithms

Genomic networks represent a complex map of molecular interactions which are descriptive of the biological processes occurring in living cells. Identifying the small over-represented circuitry patterns in these networks helps generate hypotheses about the functional basis of such complex processes. Network motif discovery is a systematic way of achieving this goal. However, a reliable network motif discovery outcome requires generating random background networks which are the result of a uniform and independent graph sampling method. To date, there has been no sound practical method to numerically evaluate whether any network motif discovery algorithm performs as intended--thus it was not possible to assess the validity of resulting network motifs. In this work, we present IndeCut, the first and only method that allows characterization of network motif finding algorithm performance on any network of interest. We demonstrate that it is critical to use IndeCut prior to running any network motif finder for two reasons. First, IndeCut estimates the minimally required number of samples that each network motif discovery tool needs in order to produce an outcome that is both reproducible and accurate. Second, IndeCut allows users to choose the most accurate network motif discovery tool for their network of interest among many available options. IndeCut is an open source software package and is available at https://github.com/megrawlab/IndeCut.

systems biology

An experimental and computational framework to build a dynamic protein atlas of human cell division

Essential biological functions, such as mitosis, require tight coordination of hundreds of proteins in space and time. Localization, timing of interactions and changes in cellular structure are all crucial to ensure correct assembly, function and regulation of protein complexes1-4. Live cell imaging can reveal protein distributions and dynamics but experimental and theoretical challenges prevented its use to produce quantitative data and a model of mitosis that comprehensively integrates information and enables analysis of the dynamic interactions between the molecular parts of the mitotic machinery within changing cellular boundaries.\n\nTo address this, we generated a 4D image data-driven, canonical model of the morphological changes during mitotic progression of human cells. We used this model to integrate dynamic 3D concentration data of many fluorescently knocked-in mitotic proteins, imaged by fluorescence correlation spectroscopy-calibrated microscopy5. The approach taken here in the context of the MitoSys consortium to generate a dynamic protein atlas of human cell division is generic. It can be applied to systematically map and mine dynamic protein localization networks that drive cell division in different cell types and can be conceptually transferred to other cellular functions.

systems biology

Multi-study inference of regulatory networks for more accurate models of gene regulation

Gene regulatory networks are composed of sub-networks that are often shared across biological processes, cell-types, and organisms. Leveraging multiple sources of information, such as publicly available gene expression datasets, could therefore be helpful when learning a network of interest. Integrating data across different studies, however, raises numerous technical concerns. Hence, a common approach in network inference, and broadly in genomics research, is to separately learn models from each dataset and combine the results. Individual models, however, often suffer from under-sampling, poor generalization and limited network recovery. In this study, we explore previous integration strategies, such as batch-correction and model ensembles, and introduce a new multitask learning approach for joint network inference across several datasets. Our method initially estimates the activities of transcription factors, and subsequently, infers the relevant network topology. As regulatory interactions are context-dependent, we estimate model coefficients as a combination of both dataset-specific and conserved components. In addition, adaptive penalties may be used to favor models that include interactions derived from multiple sources of prior knowledge including orthogonal genomics experiments. We evaluate generalization and network recovery using examples from Bacillus subtilis and Saccharomyces cerevisiae, and show that sharing information across models improves network reconstruction. Finally, we demonstrate robustness to both false positives in the prior information and heterogeneity among datasets.

systems biology