Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

A comprehensive, mechanistically detailed, and executable model of the Cell Division Cycle in Saccharomyces cerevisiae

Understanding how cellular functions emerge from the underlying molecular mechanisms is a key challenge in biology. This will require computational models, whose predictive power is expected to increase with coverage and precision of formulation. Genome-scale models revolutionised the metabolic field and made the first whole-cell model possible. However, the lack of genome-scale models of signalling networks blocks the development of eukaryotic whole-cell models. Here, we present a comprehensive mechanistic model of the molecular network that controls the cell division cycle in Saccharomyces cerevisiae. We use rxncon, the reaction-contingency language, to neutralise the scalability issues preventing formulation, visualisation and simulation of signalling networks at the genome-scale. We use parameter-free modelling to validate the network and to predict genotype-to-phenotype relationships down to residue resolution. This mechanistic genome-scale model offers a new perspective on eukaryotic cell cycle control, and opens up for similar models - and eventually whole-cell models - of human cells.

systems biology

Cell Painting, a high-content image-based assay for morphological profiling using multiplexed fluorescent dyes

In morphological profiling, quantitative data are extracted from microscopy images of cells to identify biologically relevant similarities and differences among samples based on these profiles. This protocol describes the design and execution of experiments using Cell Painting, a morphological profiling assay multiplexing six fluorescent dyes imaged in five channels, to reveal eight broadly relevant cellular components or organelles. Automated image analysis software identifies individual cells and measures ~1,500 morphological features (various measures of size, shape, texture, intensity, etc.) to produce a rich profile suitable for detecting subtle phenotypes. Profiles of cell populations treated with different experimental perturbations can be compared to suit many goals, such as identifying the phenotypic impact of chemical or genetic perturbations, grouping compounds and/or genes into functional pathways, and identifying signatures of disease. Cell culture and image acquisition takes 2 weeks; feature extraction and data analysis take an additional 1-2 weeks.

Systems Biology

GNE: A deep learning framework for gene network inference by aggregating biological information

The topological landscape of gene interaction networks provides a rich source of information for inferring functional patterns of genes or proteins. However, it is still a challenging task to aggregate heterogeneous biological information such as gene expression and gene interactions to achieve more accurate inference for prediction and discovery of new gene interactions. In particular, how to generate a unified vector representation to integrate diverse input data is a key challenge addressed here. We propose a scalable and robust deep learning framework to learn embedded representations to unify known gene interactions and gene expression for gene interaction predictions. These low-dimensional embeddings derive deeper insights into the structure of rapidly accumulating and diverse gene interaction networks and greatly simplify downstream modeling. We compare the predictive power of our deep embeddings to the strong baselines. The results suggest that our deep embeddings achieve significantly more accurate predictions. Moreover, a set of novel gene interaction predictions are validated by up-to-date literature-based database entries. GNE is freely available under the GNU General Public License and can be downloaded from Github (https://github.com/kckishan/GNE)

systems biology

A duplex MIPs-based biological-computational cell lineage discovery platform

Cell lineage analysis aims to uncover the developmental history of an organism back to its cell of origin1. Recently, novel in vivo methods and technologies utilizing genome editing enabled important insights into the cell lineages of animals2-8. In contrast, human cell lineage remains restricted to retrospective approaches, which still lack in resolution and cost-efficient solutions. Here we demonstrate a scalable platform for human cell lineage tracing based on Short Tandem Repeats (STRs) targeted by duplex Molecular Inversion Probes (MIPs). With this platform we accurately reproduced a known lineage of DU145 cell lines cells9 and reconstructed lineages of healthy and metastatic single cells from a melanoma patient. The reconstructed trees matched the anatomical and SNV references while adding further refinements. Our platform allowed to faithfully recapitulate lineages of developmental tissue formation in cells from healthy donors. In summary, our lineage discovery platform can profile informative STR somatic mutations efficiently and we provide a solid, high-resolution lineage reconstruction even in challenging low-mutation-rate healthy single cells.

systems biology

Characterizing and comparing phylogenies from their Laplacian spectrum

Phylogenetic trees are central to many areas of biology, ranging from population genetics and epidemiology to microbiology, ecology, and macroevolution. The ability to summarize properties of trees, compare different trees, and identify distinct modes of division within trees is essential to all these research areas. But despite wide-ranging applications, there currently exists no common, comprehensive framework for such analyses. Here we present a graph-theoretical approach that provides such a framework. We show how to construct the spectral density profiles of phylogenetic trees from their Laplacian graphs. Using ultrametric simulated trees as well as non-ultrametric empirical trees, we demonstrate that the spectral density successfully identifies various properties of the trees and clusters them into meaningful groups. Finally, we illustrate how the eigengap can identify modes of division within a given tree. As phylogenetic data continue to accumulate and to be integrated into various areas of the life sciences, we expect that this spectral graph-theoretical framework to phylogenetics will have powerful and long-lasting applications.

Systems Biology

Non-canonical circadian oscillations in Drosophila S2 cells drive gene-expression cycles coupled to metabolic oscillations

Circadian rhythms are cell-autonomous biological oscillations with a period of about 24 hours. Current models propose that transcriptional feedback loops are the principal mechanism for the generation of circadian oscillations. In these models, Drosophila S2 cells are generally regarded as non-rhythmic cells, as they do not express several canonical circadian components. Using an unbiased multi-omics approach, we made the surprising discovery that Drosophila S2 cells do in fact display widespread daily rhythms. Transcriptomics and proteomics analyses revealed that hundreds of genes and their products are rhythmically expressed in a 24-hour cycle. Metabolomics analyses extended these findings and illustrated that central carbon metabolism and amino acid metabolism are the main pathways regulated in a rhythmic fashion. We thus demonstrate that daily genome-wide oscillations, coupled to metabolic cycles, take place in eukaryotic cells without the contribution of known circadian regulators.

systems biology

Continuous visualization of differences between biological conditions in single-cell data

In high-dimensional single cell data, comparing changes in functional markers between conditions is typically done across manual or algorithm-derived partitions based on population-defining markers. Visualizations of these partitions is commonly done on low-dimensional embeddings (eg. t-SNE), colored by per-partition changes. Here, we provide an analysis and visualization tool that performs these comparisons across overlapping k-nearest neighbor (KNN) groupings. This allows one to color low-dimensional embeddings by marker changes without hard boundaries imposed by partitioning. We devised an objective optimization of k based on minimizing functional marker KNN imputation error. Proof-of-concept work visualized the exact location of an IL-7 responsive subset in a B cell developmental trajectory on a t-SNE map independent of clustering. Per-condition cell frequency analysis revealed that KNN is sensitive to detecting artifacts due to marker shift, and therefore can also be valuable in a quality control pipeline. Overall, we found that KNN groupings lead to useful multiple condition visualizations and efficiently extract a large amount of information from mass cytometry data. Our software is publicly available through the Bioconductor package Sconify.

systems biology

Mass-spectrometry of single mammalian cells quantifies proteome heterogeneity during cell differentiation

Cellular heterogeneity is important to biological processes, including cancer and development. However, proteome heterogeneity is largely unexplored because of the limitations of existing methods for quantifying protein levels in single cells. To alleviate these limitations, we developed Single Cell ProtEomics by Mass Spectrometry (SCoPE-MS), and validated its ability to identify distinct human cancer cell types based on their proteomes. We used SCoPE-MS to quantify over a thousand proteins in differentiating mouse embryonic stem (ES) cells. The single-cell proteomes enabled us to deconstruct cell populations and infer protein abundance relationships. Comparison between single-cell proteomes and transcriptomes indicated coordinated mRNA and protein covariation. Yet many genes exhibited functionally concerted and distinct regulatory patterns at the mRNA and the protein levels, suggesting that post-transcriptional regulatory mechanisms contribute to proteome remodeling during lineage specification, especially for developmental genes. SCoPE-MS is broadly applicable to measuring proteome configurations of single cells and linking them to functional phenotypes, such as cell type and differentiation potentials.

systems biology

Massively parallel single cell lineage tracing using CRISPR/Cas9 induced genetic scars

A key goal of developmental biology is to understand how a single cell transforms into a full-grown organism consisting of many different cell types. Single-cell RNA-sequencing (scRNA-seq) has become a widely-used method due to its ability to identify all cell types in a tissue or organ in a systematic manner 1-3. However, a major challenge is to organize the resulting taxonomy of cell types into lineage trees revealing the developmental origin of cells. Here, we present a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes generated by genome editing of transgenic reporter genes, we reconstruct developmental lineage trees in zebrafish larvae and adult fish. In future analyses, LINNAEUS (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences) can be used as a systematic approach for identifying the lineage origin of novel cell types, or of known cell types under different conditions.

systems biology

Quantitative Assessment of Protein Activity in Orphan Tissues and Single Cells Using the metaVIPER Algorithm

We and others have shown that transition and maintenance of biological states is controlled by master regulator protein, which can be inferred by interrogating tissue-specific regulatory models (interactomes) with transcriptional signatures, using the VIPER algorithm. Yet, some tissues may lack molecular profiles necessary for interactome inference (orphan tissues), or, as for single cells isolated from heterogeneous samples, their tissue context may be undetermined. To address this problem, we introduce metaVIPER, a novel algorithm designed to assess protein activity in tissue-independent by integrative analysis of multiple, non-tissue-matched interactomes. This assumes that transcriptional targets of each protein will be recapitulated by one or more available interactome. We confirmed the algorithms value in assessing protein dysregulation induced by somatic mutations, as well as in assessing protein activity in orphan tissues and, most critically, in single cells, thus allowing transformation of noisy and potentially biased RNA-Seq signatures into reproducible protein-activity signatures.

systems biology

Tn-Core: context-specific reconstruction of core metabolic models using Tn-seq data

MotivationTn-seq (transposon mutagenesis and sequencing) and constraint-based metabolic modelling represent highly complementary approaches. They can be used to probe the core genetic and metabolic networks underlying a biological process, revealing invaluable information for synthetic biology engineering of microbial cell factories. However, while algorithms exist for integration of -omics data sets with metabolic models, no method has been explicitly developed for integration of Tn-seq data with metabolic reconstructions.\n\nResultsWe report the development of Tn-Core, a Matlab toolbox designed to generate gene-centric, context-specific core reconstructions consistent with experimental Tn-seq data. Extensions of this algorithm allow: i) the generation of context-specific functional models through integration of both Tn-seq and RNA-seq data; ii) to visualize redundancy in core metabolic processes; and iii) to assist in curation of de novo draft metabolic models. The utility of Tn-Core is demonstrated primarily using a Sinorhizobium meliloti model as a case study.\n\nAvailability and implementationThe software can be downloaded from https://github.com/diCenzo-GC/Tn-Core. All results presented in this work have been obtained with Tn-Core v. 1.0.\n\nContactgeorgecolin.dicenzo@unifi.it, marco.fondi@unifi.it\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Deciphering the Dynamics of Interlocked Feedback Loops in a Model of the Mammalian Circadian Clock

Mathematical models of fundamental biological processes play an important role in consolidating theory and experiments, especially if they are systematically developed, thoroughly characterized, and well tested by experimental data. In this work, we report a detailed bifurcation analysis of a mathematical model of the mammalian circadian clock network developed by Relogio et al. [16], noteworthy for its consistency with available data. Using one- and two-parameter bifurcation diagrams, we explore how oscillations in the model depend on the expression levels of its constituent genes and the activities of their encoded proteins. These bifurcation diagrams allow us to decipher the dynamics of interlocked feedback loops, by parametric variation of genes and proteins in the model. Among other results, we find that REV-ERB, a member of a subfamily of orphan nuclear receptors, plays a critical role in the intertwined dynamics of Relogios model. The bifurcation diagrams reported here can be used for predicting how the core-clock network responds--in terms of period, amplitude and phases of oscillations--to different perturbations.

systems biology

Quantification of very low-abundant proteins in bacteria using the HaloTag and epi-fluorescence microscopy

Cell biology is increasingly dependent on quantitative methods resulting in the need for microscopic labelling technologies that are highly sensitive and specific. Whilst the use of fluorescent proteins has led to major advances, they also suffer from their relatively low brightness and photo-stability, making the detection of very low abundance proteins using fluorescent protein-based methods challenging. Here, we characterize the use of the self-labelling protein tag called HaloTag, in conjunction with an organic fluorescent dye, to label and accurately count endogenous proteins present in very low numbers (<7) in individual Escherichia coli cells. This procedure can be used to detect single molecules in fixed cells with conventional epifluorescence illumination and a standard microscope. We show that the detection efficiency of proteins labelled with the HaloTag is [&ge;]80%, which is on par or better than previous techniques. Therefore, this method offers a simple and attractive alternative to current procedures to detect low abundance molecules.

systems biology

Precise label-free quantitative proteomes in high-throughput by microLC and data-independent SWATH acquisition

While quantitative proteomics is a key technology in biological research, the routine industry and diagnostics application is so far still limited by a moderate throughput, data consistency and robustness. In part, the restrictions emerge in the proteomics dependency on nanolitre/minute flow rate chromatography that enables a high sensitivity, but is difficult to handle on large sample series, and on the stochastic nature in data-dependent acquisition strategies. We here establish and benchmark a label-free, quantitative proteomics platform that uses microlitre/minute flow rate chromatography in combination with data-independent SWATH acquisition. Being able to largely compensate for the loss of sensitivity by exploiting the analytical capacities of microflow chromatography, we show that microLC-SWATH-MS is able to precisely quantify up to 4000 proteins in an hour or less, enables the consistent processing of sample series in high-throughput, and gains quantification precisions comparable to targeted proteomic assays. MicroLC-SWATH-MS can hence routinely process hundreds to thousands of samples to systematically create precise, label free quantitative proteomes.

Systems Biology

Aggregation of biological particles under radial directional guidance.

Many biological environments display an almost radially-symmetric structure, allowing proteins, cells or animals to move in an oriented fashion. Motivated by specific examples of cell movement in tissues, pigment protein movement in pigment cells and animal movement near watering holes, we consider a class of radially-symmetric anisotropic diffusion problems, which we call the star problem. The corresponding diffusion tensor D(x) is radially symmetric with isotropic diffusion at the origin. We show that the anisotropic geometry of the environment can lead to strong aggregations and blow-up at the origin. We classify the nature of aggregation and blow-up solutions and provide corresponding numerical simulations. A surprising element of this strong aggregation mechanism is that it is entirely based on geometry and does not derive from chemotaxis, adhesion or other well known aggregating mechanisms. We use these aggregate solutions to discuss the process of pigmentation changes in animals, cancer invasion in an oriented fibrous habitat (such as collagen fibres), and sheep distributions around watering holes.\n\nJTB classification21.050, 21.160, 52.250, 71.060

systems biology

Post-inference Methods Of Prior Knowledge Incorporation In Gene Regulatory Network Inference

The regulatory interactions in a cell control cellular response to environmental and genetic perturbations. Gene regulatory network (GRN) inference from high-throughput gene expression data helps to identify unknown regulatory interactions in a cell. One of the main challenges in the GRN inference is to identify complex biological interactions from the limited information contained in the gene expression data. Using prior biological knowledge, in addition to the gene expression data, is a common method to overcome this challenge. However, only a few GRN inference methods can inherently incorporate the prior knowledge and these methods are also not among the best-ranked in benchmarking studies.\n\nWe propose to incorporate the prior knowledge after the GRN inference so that any inference method can be used. Two algorithms have been developed and tested on the well studied Escherichia coli, yeast, and realistic in silico networks. Their accuracy is higher than the best-ranking method in the latest community-wide benchmarking study. Further, one of the algorithms identifies and removes wrong interactions predicted by the inference methods. With half of the available prior knowledge of interactions, around 970 additional correct edges were obtained and 1300 wrong interactions were removed. Moreover, the limitation that only a few GRN inference methods can incorporate the prior knowledge is overcome. Therefore, a post-inference method of incorporating the prior knowledge improves accuracy, removes wrong edges, and overcomes the limitation of GRN inference methods.

systems biology

Do Cells use Passwords? Do they Encrypt Information?

Organisms must maintain proper regulation including defense and healing. Life-threatening problems may be caused by pathogens or by a multicellular organisms own cells through cancer or auto-immune disorders. Life evolved solutions to these problems that can be conceptualized through the lens of information security, which is a well-developed field in computer science. Here I argue that taking an information security view of cells is not merely semantics, but useful to explain features of signaling, regulation, and defense. An information security perspective also offers a conduit for cross-fertilization of advanced ideas from computer science, and the potential for biology to inform computer science. First, I consider whether cells use passwords, i.e., initiation sequences that are required for subsequent signals to have effects, by analyzing the concept of pioneer transcription factors in chromatin regulation and cellular reprogramming. Second, I consider whether cells may encrypt signal transduction cascades. Encryption could benefit cells by making it more difficult for pathogens or oncogenes to hijack cell networks. By using numerous molecules cells may gain a security advantage in particular against viruses, whose genome sizes are typically under selection pressure. I provide a simple conceptual argument for how cells may peform encryption through post-translational modifications, complex formation, and chromatin accessibility. I invoke information theory to provide a criterion of an entropy spike to assess whether a signaling cascade has encryption-like features. I discuss how the frequently invoked concept of context-dependency may over-simplify more advanced features of cell signaling networks, such as encryption. Therefore, by considering that biochemical networks may be even more complex than commonly realized we may be better able to understand defenses against pathogens and pathologies.

systems biology

bioLQM: a java library for the manipulation and conversion of Logical Qualitative Models of biological networks

Here we introduce bioLQM, a new Java software toolkit for the conversion, modification, and analysis of Logical Qualitative Models of biological regulatory networks, aiming to foster the development of novel complementary tools by providing core modelling operations. Based on the definition of multi-valued logical models, it implements import and export facilities, notably for the recent SBML-qual exchange format, as well as for formats used by several popular tools, facilitating the design of workflows combining these tools. Model modifications enable the definition of various perturbations, as well as model reduction, easing the analysis of large models. Another modification enables the study of multi-valued models with tools limited to the Boolean case. Finally, bioLQM provides a framework for the development of novel analysis tools. The current version implements the usual updating modes for model simulation (notably synchronous, asynchronous, and random asynchronous), as well as some static analysis features for the identification of attractors. The bioLQM software can be integrated into analysis workflows through command line and scripting interfaces. As a Java library, it further provides core data structures to the GINsim and EpiLog interactive tools, which supply graphical interfaces and additional analysis methods for cellular and multi-cellular qualitative models.

systems biology