Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Single-Cell Gene Expression Profi ling and Cell State Dynamics: Collecting Data, Correlating Data Points and Connecting the Dots

Single-cell analyses of transcript and protein expression profiles - more precisely, single-cell resolution analysis of molecular profiles of cell populations - have now entered the center stage with widespread applications of single-cell qPCR, single-cell RNA-Seq and CyTOF. These high-dimensional population snapshot techniques are complemented by low-dimensional time-resolved, microscopy-based monitoring methods. Both fronts of advance have exposed a rich heterogeneity of cell states within uniform cell populations in many biological contexts, producing a new kind of data that has stimulated a series of computational analysis methods for data visualization, dimensionality reduction, and cluster (subpopulation) identification. The next step is now to go beyond collecting data and correlating data points: to connect the dots, that is, to understand what actually underlies the identified data patterns. This entails interpreting the \"clouds of points\" in state space as a manifestation of the underlying molecular regulatory network. In that way control of cell state dynamics can be formalized as a quasi-potential landscape, as first proposed by Waddington. We summarize key methods of data acquisition and computational analysis and explain the principles that link the single-cell resolution measurements to dynamical systems theory.

Systems Biology

Stabilized Independent Component Analysis outperforms other methods in finding reproducible signals in tumoral transcriptomes

MotivationMatrix factorization methods are widely exploited in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). Applying such methods to similar independent datasets should yield reproducible inter-series outputs, though it was never demonstrated yet.\n\nResultsWe systematically test state-of-art methods of matrix factorization on several transcriptomic datasets of the same cancer type. Inspired by concepts of evolutionary bioinformatics, we design a new framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the methods reproducibility. We show that a particular protocol of application of Independent Component Analysis (ICA), accompanied by a stabilisation procedure, leads to a significant increase in the inter-series output reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other state-of-art methods. We developed a user-friendly tool BIODICA for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent publicly available transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors that can be associated to biological processes or to technological artefacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping.\n\nAvailabilityThe BIODICA tool is available from https://github.com/LabBandSB/BIODICA.\n\nContactlaura.cantini@curie.fr and andrei.zinovyev@curie.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Gene set enrichment analysis using RNA-seq-based condition-specific metabolic networks

MotivationGenome-scale metabolic networks and transcriptomic data represent complementary sources of knowledge about an organisms metabolism, yet their integration to achieve biological insight remains challenging.\n\nResultsWe investigate here condition-specific series of metabolic sub-networks constructed by successively removing genes from a comprehensive network. The optimal order of gene removal is deduced from transcriptomic data. The sub-networks are evaluated via a fitness function, which estimates their degree of alteration. We then consider how a gene set, i.e. a group of genes contributing to a common biological function, is depleted in different series of sub-networks to detect the difference between experimental conditions. The method, named metaboGSE, is validated on public data for Yarrowia lipolytica and mouse. It is shown to produce GO terms of higher specificity compared to popular gene set enrichment methods like GSEA or topGO.\n\nAvailabilityThe metaboGSE R package is available at https://cran.r-project.org/web/packages/metaboGSE.

systems biology

TuBA: Tunable Biclustering Algorithm Reveals Clinically Relevant Tumor Transcriptional Profiles in Breast Cancer

BackgroundTraditional clustering approaches for gene expression data are not well adapted to address the complexity and heterogeneity of tumors, where small sets of genes may be aberrantly co-expressed in specific subsets of tumors. Biclustering algorithms that perform local clustering on subsets of genes and conditions help address this problem. We propose a graph-based Tunable Biclustering Algorithm (TuBA) based on a novel pairwise proximity measure, examining the relationship of samples at the extremes of genes expression profiles to identify similarly altered signatures. ResultsTuBAs predictions are consistent in 3,940 Breast Invasive Carcinoma (BRCA) samples from three independent sources, employing different technologies for measuring gene expression (RNASeq and Microarray). Over 60% of biclusters identified independently in each dataset had significant agreement in their gene sets, as well as similar clinical implications. About 50% of biclusters were enriched in the ER-/HER2- (or basal-like) subtype, while more than 50% were associated with transcriptionally active copy number changes. Biclusters representing gene co-expression patterns in stromal tissue were also identified in tumor specimens. ConclusionTuBA offers a simple biclustering method that can identify biologically relevant gene co-expression signatures not captured by traditional unsupervised clustering approaches. It complements biclustering approaches that are designed to identify constant or coherent submatrices in gene expression datasets, and outperforms them in identifying a multitude of altered transcriptional profiles that are associated with observed genomic heterogeneity of diseased states in breast cancer, both within and across tumor subtypes, a promising step in understanding disease heterogeneity, and a necessary first step in individualized therapy.

systems biology

Massively parallel whole-organism lineage tracing using CRISPR/Cas9 induced genetic scars

A key goal of developmental biology is to understand how a single cell transforms into a full-grown organism consisting of many cells. Although impressive progress has been made in lineage tracing using imaging approaches, analysis of vertebrate lineage trees has mostly been limited to relatively small subsets of cells. Here we present scartrace, a strategy for massively parallel clonal analysis based on Cas9 induced genetic scars in the zebrafish.

Systems Biology

Organisation of the transcriptional regulation of genes involved in protein transactions in yeast

The topological analyses of many large-scale molecular interaction networks often provide only limited insights into network function or evolution. In this paper, we argue that the functional heterogeneity of network components, rather than network size, is the main factor limiting the utility of topological analysis of large cellular networks. We have analysed large epistatic, functional, and transcriptional regulatory networks of genes that were attributed to the following biological process groupings: protein transactions, gene expression, cell cycle, and small molecule metabolism. Control analyses were performed on networks of randomly selected genes. We identified novel biological features emerging from the analysis of functionally homogenous biological networks irrespective of their size. In particular, direct regulation by transcription as an underrepresented feature of protein transactions. The analysis also demonstrated that the regulation of the genes involved in protein transactions at the transcriptional level was orchestrated by only a small number of regulators. Quantitative proteomic analysis of nuclear- and chromatin-enriched sub-cellular fractions of yeast provided supportive evidence for the conclusions generated by network analyses.

systems biology

Systematic Elucidation and Validation of OncoProtein-Centric Molecular Interaction Maps

The largely incomplete and tissue-independent nature of cancer pathways represents a key limitation to the ability to elucidate mechanistic determinants of cancer phenotypes and to predict adaptive response to targeted therapy. To address these challenges, we propose replacing canonical cancer pathways with a more accurate, comprehensive, and context-specific architecture - dubbed a Protein-Centric molecular interaction Map (PC-Map) - representing modulators, effectors, and cognate binding-partners of any oncoprotein of interest. To reconstruct these complex molecular architectures de novo, we introduce a novel OncoSig algorithm. Validation of a lung adenocarcinoma specific (LUAD) KRAS-centric PC-Map recapitulated known KRAS biology and, more critically, identified a novel repertoire of proteins eliciting synthetic lethality in KRASG12D LUAD organoid cultures. Showing the generalizable nature of the algorithm, we elucidated PC-Maps for ten recurrently mutated oncoproteins, including KRAS, in distinct tumor contexts. This revealed a highly context-specific nature of cancers regulatory and signaling architectures to an unprecedented degree of resolution.

systems biology

INKA, an integrative data analysis pipeline for phosphoproteomic inference of active phosphokinases

Identifying (hyper)active kinases in cancer patient tumors is crucial to enable individualized treatment with specific inhibitors. Conceptually, kinase activity can be gleaned from global protein phosphorylation profiles obtained with mass spectrometry-based phosphoproteomics. A major challenge is to relate such profiles to specific kinases to identify (hyper)active kinases that may fuel growth/progression of individual tumors. Approaches have hitherto focused on phosphorylation of either kinases or their substrates. Here, we combine kinase-centric and substrate-centric information in an Integrative Inferred Kinase Activity (INKA) analysis. INKA utilizes label-free quantification of phosphopeptides derived from kinases, kinase activation loops, kinase substrates deduced from prior experimental knowledge, and kinase substrates predicted from sequence motifs, yielding a single score. This multipronged, stringent analysis enables ranking of kinase activity and visualization of kinase-substrate relation networks in a biological sample. As a proof of concept, INKA scoring of phosphoproteomic data for different oncogene-driven cancer cell lines inferred top activity of implicated driver kinases, and relevant quantitative changes upon perturbation. These analyses show the ability of INKA scoring to identify (hyper)active kinases, with potential clinical significance.

systems biology

A mechanistic link between cellular trade-offs, gene expression and growth

Intracellular processes rarely work in isolation but continually, interact with the rest of the cell. In microbes, for example, we now know that gene expression across the whole genome typically changes with growth rate. The mechanisms driving such global regulation, however, are not well understood. Here we consider three trade-offs that because of limitations in levels of cellular energy, free ribosomes, and proteins are faced by all living cells and construct a mechanistic model that comprises these trade-offs. Our model couples gene expression with growth rate and growth rate with a growing population of cells. We show that the model recovers Monod's law for the growth of microbes and two other empirical relationships connecting growth rate to the mass fraction of ribosomes. Further, we can explain growth related effects in dosage compensation by paralogs and predict host-circuit interactions in synthetic biology. Simulating competitions between strains, we find that the regulation of metabolic pathways may have evolved not to match expression of enzymes to levels of extracellular substrates in changing environments but rather to balance a trade-off between exploiting one type of nutrient over another. Although coarse-grained, the trade-offs that the model embodies are fundamental, and, as such, our modelling framework has potentially wide application, including in both biotechnology and medicine.

Systems Biology

Universal Scaling in Biochemical Networks

The application of network science to biology has advanced our understanding of the metabolism of individual organisms and the organization of ecosystems but has scarcely been applied to life at a planetary scale. To characterize planetary-scale biochemistry, we constructed biochemical networks using a global database of 28,146 annotated genomes and metagenomes, and 8,658 cataloged biochemical reactions. We uncover scaling laws governing biochemical diversity and network structure shared across levels of organization from individuals to ecosystems, to the biosphere as a whole. Comparing real biochemical networks to random chemical networks reveals the observed biological scaling is not solely a product of the biochemistry shared across life on Earth. Instead, it emerges due to how the global inventory of biochemical reactions is partitioned into individuals. We show the three domains of life are topologically distinguishable, with > 80% accuracy in predicting evolutionary domain based on biochemical network size and average topology. Taken together our results point to a deeper level of organization in biochemical networks than what has been understood so far.

systems biology

High temporal resolution of gene expression dynamics in developing mouse embryonic stem cells

Investigations of transcriptional responses during developmental transitions typically use time courses with intervals that are not commensurate with the timescales of known biological processes. Moreover, such experiments typically focus on protein-coding transcripts, ignoring the important impact of long noncoding RNAs. We evaluated coding and noncoding expression dynamics at high temporal resolution (6-hourly) in differentiating mouse embryonic stem cells and report the effects of increased temporal resolution on the characterization of the underlying molecular processes. We present a refined resolution of global transcriptional alterations, including regulatory network interactions, coding and noncoding gene expression changes as well as alternative splicing events, many of which cannot be resolved by existing coarse developmental time-{-}-courses. We describe novel short lived and cycling patterns of gene expression and temporally dissect ordered gene expression at bidirectional promoters and responses to transcription factors. These findings demonstrate the importance of temporal resolution for understanding gene interactions in mammalian systems.\n\nLinks to dataData has been deposited into GEO: The Reviewer access link is: http://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?token=cnglummejbkltyj@acc=GSE75028

systems biology

Bifurcation analysis of metabolic pathways: an illustration from yeast glycolysis

In microorganisms such as bacteria or yeasts, metabolic rates are tightly coupled to growth rate, and therefore to fitness. Although the topology of central pathways are largely conserved across organisms, the enzyme kinetics and their parameters generally vary. This prevents us to understand and predict (changes in) metabolic dynamics. The analytical treatment of metabolic pathways is generally restricted to small models, containing maybe two to four equations. Since such small core models involve much coarse graining, their biological interpretation is often hampered. In this paper we aim to bridge the gap between analytical, more in-depth treatment of small core models and biologically more realistic and detailed models by developing new methods. We illustrate these methods for a model of glycolysis in Saccharo-myces cerevisiae yeast, arguably the best characterised metabolic pathway in the literature. The model is more involved than in previous studies, and involves both ATP/ADP and NADH/NAD householding.\n\nA detailed analysis of the steady state equations sheds new light on two recently studied biological phenomena in yeast glycolysis: whether it is to be expected that fructose-1,6-biphosphate (FBP) parameterises all steady states, and the occurrence of bistability between a regular steady state and imbalanced steady state in which glycolytic intermediates keep accumulating.\n\nThis work shows that the special structure of metabolic pathways does allow for more in-depth bifurcation analyses than is currently the norm. We especially emphasise which of the techniques developed here scale to larger pathways, and which do not.

systems biology

Analytic framework for stochastic binary biological switches

We propose an analytic solution for the stochastic dynamics of a binary biological switch, defined as a DNA unit with two mutually exclusive configurations, each one triggering the expression of a different gene. Such a device could be used as a memory unit for biological computing systems designed to operate in noisy environments. We discuss a recent implementation of an exclusive switch in living cells, the recombinase addressable data (RAD) module. In order to understand the behavior of a RAD module we compute the exact time dependent distributions of the two expressed genes starting in one state and evolving to another asymptotic state. We consider two operating regimes of the RAD module: fast and slow stochastic switching. The fast regime is \"aggregative\" and produces unimodal distributions, whereas the slow regime is \"separative\" and produces bimodal distributions. Both regimes can serve to prepare pure memory states when all cells are expressing the same gene. The slow regime can also separate mixed states by producing two sub-populations each one expressing a different gene. Our model provides a simplified, general phenomenological framework for studying biological memory devices and our analytic solution can be further used to clarify theoretical concepts in bio-computation and for optimal design in synthetic biology.

Synthetic Biology

Untargeted Metabolomics Suffers From Incomplete Data Analysis

IntroductionUntargeted metabolomics is a powerful tool for biological discoveries. Significant advances in computational approaches to analyzing the complex raw data have been made, yet it is not clear how exhaustive and reliable are the data analysis results.\n\nObjectivesAssessment of the quality of data analysis results in untargeted metabolomics.\n\nMethodsFive published untargeted metabolomics studies acquired using instruments from different manufacturers were reanalyzed.\n\nResultsOmissions of at least 50 relevant compounds from original results as well as examples of representative mistakes are reported for each study.\n\nConclusionIncomplete data analysis shows unexplored potential of current and legacy data.

systems biology

Multi-generational silencing dynamics control cell aging

Cellular aging plays an important role in many diseases, such as cancers, metabolic syndromes and neurodegenerative disorders. There has been steady progress in identifying aging-related factors such as reactive oxygen species and genomic instability, yet an emerging challenge is to reconcile the contributions of these factors with the fact that genetically identical cells can age at significantly different rates. Such complexity requires single-cell analyses designed to unravel the interplay of aging dynamics and cell-to-cell variability. Here we use novel microfluidic technologies to track the replicative aging of single yeast cells and reveal that the temporal patterns of heterochromatin silencing loss regulate cellular lifespan. We found that cells show sporadic waves of silencing loss in the heterochromatic ribosomal DNA (rDNA) during the early phases of aging, followed by sustained loss of silencing preceding cell death. Isogenic cells have different lengths of the early intermittent silencing phase that largely determine their final lifespans. Combining computational modeling and experimental approaches, we found that the intermittent silencing dynamics is important for longevity and is dependent on the conserved Sir2 deacetylase, whereas either sustained silencing or sustained loss of silencing shortens lifespan. These findings reveal, for the first time, that the temporal patterns of a key molecular process can directly influence cellular aging and thus could provide guidance for the design of temporally controlled strategies to extend lifespan.\n\nSignificanceAging is an inevitable consequence of living, and with it comes increased morbidity and mortality. Novel approaches to mitigating age-related chronic diseases demand a better understanding of the biology of aging. Studies in model organisms have identified many conserved molecular factors that influence aging. The emerging challenge is to understand how these factors interact and change dynamically to drive aging. Using multidisciplinary technologies, we have revealed a sirtuin-dependent intermittent pattern of chromatin silencing during yeast aging that is crucial for longevity. Our findings highlight the important role of silencing dynamics in aging, which deserves careful consideration when designing schemes to delay or reverse aging by modulating sirtuins and silencing.

systems biology

Identification of Cancer-associated Metabolic Vulnerabilities by Modeling Multi-objective Optimality in Metabolism

Computational modeling of the genome-wide metabolic network is essential for designing new therapeutics targeting cancer-associated metabolic disorder, which is a hallmark of human malignancies. However, previous studies generally assumed that metabolic fluxes of cancer cells are subjected to the maximization of biomass production, despite the wide existence of trade-offs among multiple metabolic objectives. To address this issue, we developed a multi-objective model of cancer metabolism with algorithms depicting approximate Pareto surfaces and incorporating multiple omics datasets. To validate this approach, we built individualized models for NCI-60 cancer cell lines, and accurately predicted cell growth rates and other biological consequences of metabolic perturbations in these cells. By analyzing the landscape of approximate Pareto surface, we identified a list of metabolic targets essential for cancer cell proliferation and the Warburg effect, and further demonstrated their close association with cancer patient survival. Finally, metabolic targets predicted to be essential for tumor progression were validated by cell-based experiments, confirming this multi-objective modelling method as a novel and effective strategy to identify cancer-associated metabolic vulnerabilities.

systems biology

Genome-wide post-transcriptional dysregulation by microRNAs in human asthma as revealed by Frac-seq

MicroRNAs are small non-coding RNAs that inhibit gene expression post-transcriptionally, implicated in virtually all biological processes. Although the effect of individual microRNAs is generally studied, the genome-wide role of multiple microRNAs is less investigated. We assessed paired genome-wide expression of microRNAs with total (cytoplasmic) and translational (polyribosome-bound) mRNA levels employing Frac-seq in human primary bronchoepithelium from healthy controls and severe asthmatics. Severe asthma is a chronic inflammatory disease of the airways characterized by poor response to therapy. We found genes (=all isoforms of a gene) and mRNA isoforms differentially expressed in asthma, with novel inflammatory and structural mechanisms disclosed solely by polyribosome-bound mRNAs. Gene expression (=all isoforms of a gene) and mRNA expression analysis revealed different molecular candidates and biological pathways, with differentially expressed polyribosome-bound and total mRNAs also showing little overlap. We reveal a hub of six dysregulated microRNAs accounting for [~]90% of all microRNA targeting, displaying preference for polyribosome-bound mRNAs. Transfection of this hub in healthy cells mimicked asthma characteristics. Our work demonstrates extensive post-transcriptional gene dysregulation in asthma, where microRNAs play a central role, illustrating the feasibility and importance of assessing post-transcriptional gene expression when investigating human disease.

systems biology

The architecture of the human RNA-binding protein regulatory network

RNA-binding proteins (RBPs) are key players of post-transcriptional regulation of gene expression. These proteins influence both cellular physiology and pathology by regulating processes ranging from splicing and polyadenylation to mRNA localization, stability, and translation. To fine-tune the outcome of their regulatory action, RBPs rely on an intricate web of competitive and cooperative interactions. Several studies have described individual interactions of RBPs with RBP mRNAs, suggestive of a RBP-RBP regulatory structure. Here we present the first systematic investigation of this structure, based on a network including almost fifty thousand experimentally determined interactions between RBPs and bound RBP mRNAs.\n\nOur analysis identified two features defining the structure of the RBP-RBP regulatory network. What we call \"RBP clusters\" are groups of densely interconnected RBPs which co-regulate their targets, suggesting a tight control of cooperative and competitive behaviors. \"RBP chains\", instead, are hierarchical structures driven by evolutionarily ancient RBPs, which connect the RBP clusters and could in this way provide the flexibility to coordinate the tuning of a broad set of biological processes.\n\nThe combination of these two features suggests that RBP chains may use the modulation of their RBP targets to coordinately control the different cell programs controlled by the RBP clusters. Under this island-hopping model, the regulatory signal flowing through the chains hops from one RBP cluster to another, implementing elaborate regulatory plans to impact cellular phenotypes. This work thus establishes RBP-RBP interactions as a backbone driving post-transcriptional regulation of gene expression to allow the fine-grained control of RBPs and their targets.

Systems Biology