Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Tracking stem cell differentiation without biomarkers using pattern recognition and phase contrast imaging

Bio-image informatics is the systematic application of image analysis algorithms to large image datasets to provide an objective method for accurately and consistently scoring image data. Within this field, pattern recognition (PR) is a form of supervised machine learning where the computer identifies relevant patterns in groups (classes) of images after being trained on examples. Rather than segmentation, image-specific algorithms or adjustable parameter sets, PR relies on extracting a common set of image descriptors (features) from the entire image to determine similarities and differences between image classes.\n\nGross morphology can be the only available description of biological systems prior to their molecular characterization, but these descriptions can be subjective and qualitative. In principle, generalized PR can provide an objective and quantitative characterization of gross morphology, thus providing a means of computationally defining morphological biomarkers. In this study, we investigated the potential of a pattern recognition approach to a problem traditionally addressed using genetic or biochemical biomarkers. Often these molecular biomarkers are unavailable for investigating biological processes that are not well characterized, such as the initial steps of stem cell differentiation.\n\nHere we use a general contrast technique combined with generalized PR software to detect subtle differences in cellular morphology present in early differentiation events in murine embryonic stem cells (mESC) induced to differentiate by the overexpression of selected transcription factors. Without the use of reporters, or a priori knowledge of the relevant morphological characteristics, we identified the earliest differentiation event (3 days), reproducibly distinguished eight morphological trajectories, and correlated morphological trajectories of 40 mESC clones with previous micro-array data. Interestingly, the six transcription factors that caused the greatest morphological divergence from an ESC-like state were previously shown by expression profiling to have the greatest influence on the expression of downstream genes.

cell biology

In vivo generation of DNA sequence diversity for cellular barcoding

Heterogeneity is a ubiquitous feature of biological systems. A complete understanding of such systems requires a method for uniquely identifying and tracking individual components and their interactions with each other. We have developed a novel method of uniquely tagging individual cells in vivo with a genetic \"barcode\" that can be recovered by DNA sequencing. We demonstrate the feasibility of this technique in bacterial cells. This method should prove useful in tracking interactions of cells within a network, and/or heterogeneity within complex biological samples.

Synthetic Biology

A data-driven approach to characterising intron signal in RNA-seq data

RNA-seq datasets can contain millions of intron reads per sequenced library that are typically removed from downstream analysis. Only reads overlapping annotated exons are considered to be informative since mature mRNA is assumed to be the major component sequenced, especially when examining poly(A) RNA samples. In this paper, we demonstrate that intron reads are informative and that pre-mRNA is the major source of intron signal. Making use of pre-mRNA signal, our index method combines differential expression analyses from intron and exon counts to categorise changes observed in each count set, giving additional genes with evidence of transcriptional changes when compared to a classic approach. Considering the importance of intron retention in some biological systems, another novel method, superintronic, looks for evidence of intron retention after accounting for the presence of pre-mRNA signal. The results presented here overcomes deficiencies and biases in previous works related to intron reads by exploring multiple sources for intron reads simultaneously using a data-driven approach, and provides a broad overview into how intron reads can be utilised in relation to multiple aspects of transcriptional biology.

bioinformatics

Pathogen Population Structure Can Explain Hospital Outbreaks

ObjectiveTo analyze Hospital Acquired Infection (HAI) outbreaks using microbial population biology dynamics in order to understand outbreaks as a biological system.\n\nDesignComputational modeling study.\n\nMethodsThe majority of HAI transmission models describe dynamics on the level of the host rather than on the level of the pathogens themselves. Accordingly, epidemiologists often cannot complete transmission chains without direct evidence of either host-host contact or large reservoir populations. Here, we propose an ecology-based model to explain the transmission of pathogens in hospitals. The model is based upon metapopulation biology, which describes a group of interacting localized populations and island biogeography, which provides a basis for how pathogens may be moving between locales. Computational simulation trials are used to assess the applicability of the model.\n\nResultsResults indicate that pathogens survive for extended periods without the need for large reservoirs by living in localized ephemeral populations while continuously transmitting pathogens to new seed populations. Computational simulations show small populations spending significant portions of time at sizes too small to be detected by most surveillance protocols. The number and type of these ephemeral populations enable the overall pathogen population to be sustained.\n\nConclusionsBy modeling hospital pathogens as a metapopulation, observations characteristic of hospital acquired infection outbreaks for which there has previously been no sufficient biological explanation, including how and why empirically successful interventions work, can now be accounted for using population dynamic hypotheses. Epidemiological links between temporally isolated outbreaks are explained via pathogen population dynamics and potential outbreak intervention targets are identified.

epidemiology

A Synthetic Microbial Operational Amplifier

Synthetic biology has created oscillators, latches, logic gates, logarithmically linear circuits, and load drivers that have electronic analogs in living cells. The ubiquitous operational amplifier, which allows circuits to operate robustly and precisely has not been built with bio-molecular parts. As in electronics, a biological operational-amplifier could greatly improve the predictability of circuits despite noise and variability, a problem that all cellular circuits face. Here, we show how to create a synthetic 3-stage inducer-input operational amplifier with a differential transcription-factor stage, a CRISPR-based push-pull stage, and an enzymatic output stage with just 5 proteins including dCas9. Our Bio-OpAmp expands the toolkit of fundamental circuits available to bioengineers or biologists, and may shed insight into biological systems that require robust and precise molecular homeostasis and regulation.\n\nOne Sentence SummaryA synthetic bio-molecular operational amplifier that can enable robust, precise, and programmable homeostasis and regulation in living cells with just 5 protein parts is described.

synthetic biology

An amplicon-based sequencing framework for accurately measuring intrahost virus diversity using PrimalSeq and iVar

How viruses evolve within hosts can dictate infection outcomes; however, reconstructing this process is challenging. We evaluated our multiplexed amplicon approach - PrimalSeq - to demonstrate how virus concentration, sequencing coverage, primer mismatches, and replicates influence the accuracy of measuring intrahost virus diversity. We developed an experimental protocol and computational tool (iVar) for using PrimalSeq to measure virus diversity using Illumina and compared the results to Oxford Nanopore sequencing. We demonstrate the utility of PrimalSeq by measuring Zika and West Nile virus diversity from varied sample types and show that the accumulation of genetic diversity is influenced by experimental and biological systems.

evolutionary biology

PathMe: Merging and exploring mechanistic pathway knowledge

BackgroundThe complexity of representing biological systems is compounded by an ever-expanding body of knowledge emerging from multi-omics experiments. A number of pathway databases have facilitated pathway-centric approaches that assist in the interpretation of molecular signatures yielded by these experiments. However, the lack of interoperability between pathway databases has hindered the ability to harmonize these resources and to exploit their consolidated knowledge. Such a unification of pathway knowledge is imperative in enhancing the comprehension and modeling of biological abstractions.\n\nResultsHere, we present PathMe, a Python package that transforms pathway knowledge from three major pathway databases into a unified abstraction using Biological Expression Language as the pivotal, integrative schema. PathMe is complemented by a novel web application (freely available at https://pathme.scai.fraunhofer.de/) which allows users to comprehensively explore pathway crosstalks and compare areas of consensus and discrepancies.\n\nConclusionsThis work has harmonized three major pathway databases and transformed them into a unified schema in order to gain a holistic picture of pathway knowledge. We demonstrate the utility of the PathMe framework in: i) integrating pathway landscapes at the database level, ii) comparing the degree of consensus at the pathway level, and iii) exploring pathway crosstalk and investigating consensus at the molecular level.

bioinformatics

ComPath: An ecosystem for exploring, analyzing, and curating pathway databases

Although pathways are widely used for the analysis and representation of biological systems, their lack of clear boundaries, their dispersion across numerous databases, and the lack of interoperability impedes the evaluation of the coverage, agreements, and discrepancies between them. Here, we present ComPath, an ecosystem that supports curation of pathway mappings between databases and fosters the exploration of pathway knowledge through several novel visualizations. We have curated mappings between three of the major pathway databases and present a case study focusing on Parkinsons disease that illustrates how ComPath can generate new biological insights by identifying pathway modules, clusters, and cross-talks with these mappings. The ComPath source code and resources are available at https://github.com/ComPath and the web application can be accessed at http://compath.scai.fraunhofer.de/.

bioinformatics

NetREX: Network Rewiring using EXpression - Towards Context Specific Regulatory Networks

Understanding gene regulation is a fundamental step towards understanding of how cells function and respond to environmental cues and perturbations. An important step in this direction is the ability to infer the transcription factor (TF)-gene regulatory network (GRN). However gene regulatory networks are typically constructed disregarding the fact that regulatory programs are conditioned on tissue type, developmental stage, sex, and other factors. Due to lack of the biological context specificity, these context-agnostic networks may not provide insight for revealing the precise actions of genes for a specific biological system under concern. Collecting multitude of features required for a reliable construction of GRNs such as physical features (TF binding, chromatin accessibility) and functional features (correlation of expression or chromatin patterns) for every context of interest is costly. Therefore we need methods that is able to utilize the knowledge about a context-agnostic network (or a network constructed in a related context) for construction of a context specific regulatory network.\n\nTo address this challenge we developed a computational approach that utilizes expression data obtained in a specific biological context such as a particular development stage, sex, tissue type and a GRN constructed in a different but related context (alternatively an incomplete or a noisy network for the same context) to construct a context specific GRN. Our method, NetREX, is inspired by network component analysis (NCA) that estimates TF activities and their influences on target genes given predetermined topology of a TF-gene network. To predict a network under a different condition, NetREX removes the restriction that the topology of the TF-gene network is fixed and allows for adding and removing edges to that network. To solve the corresponding optimization problem, which is non-convex and non-smooth, we provide a general mathematical framework allowing use of the recently proposed Proximal Alternative Linearized Maximization technique and prove that our formulation has the properties required for convergence.\n\nWe tested our NetREX on simulated data and subsequently applied it to gene expression data in adult females from 99 hemizygotic lines of the Drosophila deletion (DrosDel) panel. The networks predicted by NetREX showed higher biological consistency than alternative approaches. In addition, we used the list of recently identified targets of the Doublesex (DSX) transcription factor to demonstrate the predictive power of our method.

bioinformatics

The Evolutionary Landscape of Pan-Cancer Drives Clinical Aggression

Although cancer mechanisms differ from occurrence and development, some of them have similar oncogenesis, which leads to similar clinical phenotypes. Most existing genotyping studies look at \"omics\" data, but intentionally or unintentionally avoided that cancer is a time-dependent evolutionary process, biologically represented by the time evolution of tumor clones. We used the Bayesian mutation landscape approach to reconstruct the evolutionary process of cancer by acquiring somatic mutation data consisting of 21 cancer types. Four representative evolution patterns of pan-cancer have been discovered: trees, chaos, biconvex, and Cambrian, and a strong correlation between these four evolutionary patterns and clinical aggressivity. We further explained the characteristics of the corresponding biological systems in the evolution of pan cancer by analyzing the function of differentially expressed protein-protein interaction networks. Our results explained the difference in clinical aggressivity between cancer evolution patterns from the evolution of tumor clones and exposed the functional mechanism behind.

cancer biology

Evolutionarily informed deep learning methods: Predicting transcript abundance from DNA sequence

Deep learning methodologies have revolutionized prediction in many fields, and show potential to do the same in molecular biology and genetics. However, applying these methods in their current forms ignores evolutionary dependencies within biological systems and can result in false positives and spurious conclusions. We developed two novel approaches that account for evolutionary relatedness in machine learning models: 1) gene-family guided splitting, and 2) ortholog contrasts. The first approach accounts for evolution by constraining the models training and testing sets to include different gene families. The second, uses evolutionarily informed comparisons between orthologous genes to both control for and leverage evolutionary divergence during the training process. The two approaches were explored and validated within the context of mRNA expression level prediction, and have prediction auROC values ranging from 0.72 to 0.94. Model weight inspections showed biologically interpretable patterns, resulting in the novel hypothesis that the 3 UTR is more important for fine tuning mRNA abundance levels while the 5 UTR is more important for large scale changes.

molecular biology

Integrative genomics study of microglial transcriptome reveals effect of DLG4 (PSD95) on white matter in preterm infants.

Preterm birth places newborn infants in an adverse environment that leads to brain injury linked to neuroinflammation. To characterise this pathology, we present a translational bioinformatics investigation, with integration of human and mouse molecular and neuroimaging datasets to provide a deeper understanding of the role of microglia in preterm white matter damage. We examined preterm neuroinflammation in a mouse model of encephalopathy of prematurity induced by IL1B exposure, carrying out a gene network analysis of the cell-specific transcriptomic response to injury, which we extended to analysis of protein-protein interactions, transcription factors, and human brain gene expression, including translation to preterm infants by means of imaging-genetics approaches in the brain. We identified the endogenous synthesis of DLG4 (PSD95) protein by microglia in mouse and human, modulated by inflammation and development. Systemic genetic variation in DLG4 was associated with structural features in the preterm infant brain, suggesting that genetic variation in DLG4 may also impact white matter development and inter-individual susceptibility to injury.\n\nPreterm birth accounts for 11% of all births 1, and is the leading global cause of deaths under 5 years of age 2. Over 30% of survivors experience motor and/or cognitive problems from birth 3, 4, which last into adulthood 5. These problems include a 3-8 fold increased risk of symptoms and disorders associated with anxiety, inattention and social and communication problems compared to term-born infants 6. Prematurity is associated with a 4-12 fold increase in the prevalence of Autism Spectrum Disorders (ASD) compared to the general population 7, as well as a risk ratio of 7.4 for bipolar affective disorder among infants born below 32 weeks of gestation 8.\n\nThe characteristic brain injury observed in contemporary cohorts of preterm born infants includes changes to the grey and white matter tissues, that specifically include oligodendrocyte maturation arrest, hypomyelination and cortical changes visualised as decreases in fractional anisotropy 9-13. Exposure of the fetus and postnatal infant to systemic inflammation is an important contributing factor to brain injury in preterm born infants 12, 14, 15, and the persistence of inflammation is associated with poorer neurological outcome 16. Sources of systemic inflammation include maternal/fetal infections such as chorioamnionitis (which it is estimated affects a large number of women at a sub-clinical level), with the effect of systemic inflammation in the brain being mediated predominantly by the microglial response 17.\n\nMicroglia are unique yolk-sac derived resident phagocytes of the brain 18, 19, found preferentially within the developing white matter as a matter of normal developmental migration 12. Microglial products associated with white matter injury include pro-inflammatory cytokines, such as interleukin-1{beta} (IL1B) and tumour necrosis factor (TNF-)20, which can lead to a sub-clinical inflammatory situation associated with unfavourable outcomes 21. In addition to being key effector cells in brain inflammation, they are critical for normal brain development in processes such as axonal growth and synapse formation 22, 23. The role of microglia in neuroinflammation is dynamic and complex, reflected in their mutable phenotypes including both pro-inflammatory and restorative functions 24. Despite their important neurobiological role, the time course and nature of the microglial responses in preterm birth are currently largely unknown, and the interplay of inflammatory and developmental processes is also unclear. We, and others, believe that a better understanding of the molecular mechanisms underlying microglial function could harness their beneficial effects and mitigate the brain injury of prematurity and other states of brain inflammation25, 26\n\nA clinically relevant experimental mouse model of IL1B-induced systemic inflammation has been developed to study the changes occurring in the preterm human brain 27, 28. This model recapitulates the hallmarks of encephalopathy of prematurity including oligodendrocyte maturation delay with consequent dysmyelination, associated magnetic resonance imaging (MRI) phenotypes and behavioural deficits. Here, we take advantage of this model system to characterise the molecular underpinnings of the microglial response to IL1B-driven systemic inflammation and investigate its role in concurrent development.\n\nIn preterm infants MRI is used extensively to provide in-vivo correlates of white and grey matter pathology, allowing clinical assessment and prognostication. Diffusion MRI (d-MRI) measures the displacement of water molecules in the brain, and provides insight into the underlying tissue structure. Various d-MRI measures of white matter have been associated with developmental outcome in children born preterm 29-32, with up to 60% of inter-individual variability in structural and functional features attributable to genetic factors 33, 34. White matter abnormalities are linked to associated grey matter changes at both the imaging and cellular level 10, 35, 36, with functional and structural consequences lasting into adulthood 37, 38. Tract Based Statistics (TBSS) allows quantitative whole-brain white matter analysis of d-MRI data at the voxel level while avoiding problems due to contamination by signals arising from grey matter 39. This permits voxel-wise statistical testing and inferences to be made about group differences or associations with greater statistical power. TBSS has been shown to be an effective tool for studying white matter development and injury in the preterm brain 40, providing a macroscopic in vivo quantitative measure of white matter integrity that is associated with cognitive, fine motor, and gross motor outcome 11, 41, 42.\n\nIn this work we take a translational systems biology approach to investigate the role of microglia in preterm neuroinflammation and brain injury. We integrate microglial cell-type specific data from a mouse model of perinatal neuroinflammatory brain injury with experimental ex vivo and in vitro validation, translation to the human brain across the lifespan including analysis of human microglia, and assessment of the impact of genetic variation on structure of the preterm brain. We add to the understanding of the neurobiology of prematurity by: a) revealing the endogenous expression of DLG4 (PSD95) by microglia in early development, which is modulated by developmental stage and inflammation; and b) finding an association between systemic genetic variability in DLG4 and white matter structure in the preterm neonatal brain.

genomics

DataPackageR: Reproducible data preprocessing, standardization and sharing using R/Bioconductor for collaborative data analysis.

A central tenet of reproducible research is that scientific results are published along with the underlying data and software code necessary to reproduce and verify the findings. A host of tools and software have been released that facilitate such work-flows and scientific journals have increasingly demanded that code and primary data be made available with publications. There has been little practical advice on implementing reproducible research work-flows for large omics or systems biology data sets used by teams of analysts working in collaboration. In such instances it is important to ensure all analysts use the same version of a data set for their analyses. Yet, instantiating relational databases and standard operating procedures can be unwieldy, with high \"startup\" costs and poor adherence to procedures when they deviate substantially from an analysts usual work-flow. Ideally a reproducible research work-flow should fit naturally into an individuals existing work-flow, with minimal disruption. Here, we provide an overview of how we have leveraged popular open source tools, including Bioconductor, Rmarkdown, git version control, R, and specifically Rs package system combined with a new tool DataPackageR, to implement a lightweight reproducible research work-flow for preprocessing large data sets, suitable for sharing among small-to-medium sized teams of computational scientists. Our primary contribution is the DataPackageR tool, which decouples time-consuming data processing from data analysis while leaving a traceable record of how raw data is processed into analysis-ready data sets. The software ensures packaged data objects are properly documented and performs checksum verification of these along with basic package version management, and importantly, leaves a record of data processing code in the form of package vignettes. Our group has implemented this work-flow to manage, analyze and report on pre-clinical immunological trial data from multi-center, multi-assay studies for the past three years.

bioinformatics

Gsmodutils: A python based framework for test-driven genome scale metabolic model development

MotivationGenome scale metabolic models (GSMMs) are increasingly important for systems biology and metabolic engineering research as they are capable of simulating complex steady-state behaviour. Constraints based models of this form can include thousands of reactions and metabolites, with many crucial pathways that only become activated in specific simulation settings. However, despite their widespread use, power and the availability of tools to aid with the construction and analysis of large scale models, little methodology is suggested for the continued management of curated large scale models. For example, when genome annotations are updated or new understanding regarding behaviour of is discovered, models often need to be altered to reflect this. This is quickly becoming an issue for industrial systems and synthetic biotechnology applications, which require good quality reusable models integral to the design, build and test cycle.\n\nResultsAs part of an ongoing effort to improve genome scale metabolic analysis, we have developed a test-driven development methodology for the continuous integration of validation data from different sources. Contributing to the open source technology based around COBRApy, we have developed the gsmodutils modelling framework placing an emphasis on test-driven design of models through defined test cases. Crucially, different conditions are configurable allowing users to examine how different designs or curation impact a wide range of system behaviours, minimising error between model versions.\n\nAvailabilityThe software framework described within this paper is open source and freely available from http://github.com/SBRCNottingham/gsmodutils

bioinformatics

Information theory and the phenotypic complexity of evolutionary adaptations and innovations

Two main lines of research link information theory to evolutionary biology. The first focuses on organismal phenotypes, and on the information that organisms acquire about their environment. The second connects information-theoretic concepts to genotypic change. The genotypic and phenotypic level can be linked by experimental high-throughput genotyping and computational models of genotype-phenotype relationships. I here use a simple information-theoretic framework to compute a phenotypes information content (its phenotypic complexity), and the information gain or change that comes with a new phenotype. I apply this framework to experimental data on DNA-binding phenotypes of multiple transcription factors. Low phenotypic complexity is associated with a biological systems ability to discover novel phenotypes in evolution. I show that DNA duplications lower phenotypic complexity, which illustrates how information theory can help explain why gene duplications accelerate evolutionary adaptation. I also demonstrate that with the right experimental design, sequencing data can be used to infer the information gain associated with novel evolutionary adaptations, for example in laboratory evolution experiments. Information theory can help quantify the evolutionary progress embodied in the discovery of novel adaptive phenotypes.

Evolutionary Biology

A novel mathematical method for disclosing oscillations ingene transcription: a comparative study

Circadian rhythmicity, the 24-hour cycle responsive to light and dark, is determined by periodic oscillations in gene transcription. This phenomenon has broad ramifications in physiologic function. Recent work has disclosed more cycles in gene transcription, and to the uncovering of these we apply a novel signal processing methodology known as the pencil method and compare it to conventional parametric, nonparametric, and statistical methods. Methods: In order to assess periodicity of gene expression over time, we analyzed a database derived from livers of mice entrained to a 12-hour light/12-hour dark cycle. We also analyzed artificially generated signals to identify differences between the pencil decomposition and other alternative methods.\n\nResultsThe pencil decomposition revealed hitherto-unsuspected oscillations in gene transcription with 12-hour periodicity. The pencil method was robust in detecting the 24-hour circadian cycle that was known to exist, as well as confirming the existence of shorter-period oscillations. A key consequence of this approach is that orthogonality of the different oscillatory components can be demonstrated. thus indicating a biological independence of these oscillations, that has been subsequently confirmed empirically by knocking out the gene responsible for the 24-hour clock.\n\nConclusionSystem identification techniques can be applied to biological systems and can uncover important characteristics that may elude visual inspection of the data. Significance: The pencil method provides new insights on the essence of gene expression and discloses a wide variety of oscillations in addition to the well-studied circadian pattern. This insight opens the door to the study of novel mechanisms by which oscillatory gene expression signals exert their regulatory effect on cells to influence human diseases.

cell biology

Harnessing Escherichia coli motility to engineer bacterial Voronoi patterns

Cell motility drives spatial pattern formation across diverse biological systems. Here, we engineer Escherichia coli motility in semi-solid agar to control Voronoi patterns in two and three dimensions, partitioning space into regions closest to their respective inoculation seeds. Consistent with our reaction-diffusion model, we observed that collisions between expansion fronts generate either biomass depletion (''gaps'') or accumulation (''anti-gaps''), governed by the relative diffusion rates of bacteria and nutrients. By engineering strains with distinct expansion rates and tuneable motility, and by integrating these experimental data into a dynamic Voronoi model, we achieved precise control over pattern geometry. This enabled the generation of gaps with varying widths, curved boundaries, asymmetric structures, seedless regions, and complex composite patterns. Together, these findings establish bacterial Voronoi patterns as a programmable platform for engineering multicellular spatial organization, with potential applications in synthetic biology and materials science.

synthetic biology

Combined aptamer and transcriptome sequencing of single cells

The transcriptome and proteome encode distinct information that is important for characterizing heterogeneous biological systems. We demonstrate a method to simultaneously characterize the transcriptomes and proteomes of single cells at high throughput using aptamer probes and droplet-based single cell sequencing. With our method, we differentiate distinct cell types based on aptamer surface binding and gene expression patterns. Aptamers provide advantages over antibodies for single cell protein characterization, including rapid, in vitro, and high-purity generation via SELEX, and the ability to amplify and detect them with PCR and sequencing.

cell biology