Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

A graph-based algorithm for RNA-seq data normalization

The use of RNA-sequencing has garnered much attention in the recent years for characterizing and understanding various biological systems. However, it remains a major challenge to gain insights from a large number of RNA-seq experiments collectively, due to the normalization problem. Current normalization methods are based on assumptions that fail to hold when RNA-seq profiles become more abundant and heterogeneous. We present a normalization procedure that does not rely on these assumptions, or on prior knowledge about the reference transcripts in those conditions. This algorithm is based on a graph constructed from intrinsic correlations among RNA-seq transcripts and seeks to identify a set of densely connected vertices as references. Application of this algorithm on our benchmark data showed that it can recover the reference transcripts with high precision, thus resulting in high-quality normalization. As demonstrated on a real data set, this algorithm gives good results and is efficient enough to be applicable to real-life data.\n\n2012 ACM Subject ClassificationApplied computing [->] Computational transcriptomics, Applied computing [->] Bioinformatics\n\nDigital Object Identifier10.4230/LIPIcs.WABI.2018.xxx\n\nFundingThis material was based on research supported by the National Heart, Lung, and Blood Institute (NHLBI)-NIH sponsored Programs of Excellence in Glycosciences [grant number HL107152 to B.K.], and partially by NSF [CAREER grant 1350344 to M.M.]. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon.

bioinformatics

Underestimation of carbohydrates by sugar alcohols in classical anthrone-based colorimetric techniques compromises insect metabolic and energetic studies

Physiologically based metabolic studies usually search for easy, sensitive, and cheap techniques to rapidly estimate biological parameters such as nutrient content. Colorimetric methods to estimate carbohydrates have been extensively used (over 120,000 references). However, sugar alcohols are underestimated under conventionally used analytical conditions, in particular if using the popular van Handel method. This may lead to misinterpretations of sugar implications in biological systems. We determined the anthrone reaction with various sugar alcohols and non-alcohols individually under standard conditions (Van Handel 1985). We then manipulated the proportion of either sugar alcohols or non-alcohols in three different sugar mixtures in order to estimate the impact on the total sugar estimation. In the case of a mixture with over 50% of sugar alcohols, total sugars are underestimated by 50% when using glucose as standard.

ecology

Phenotypic plasticity facilitates alterations in life-history strategies under combinations of environmental stresses

Plants developed various reversible and non-reversible acclimation mechanisms to cope with the multifaceted nature of abiotic stress combinations. We hypothesized that in order to endure these stress combinations, plants elicit distinctive acclimation strategies through specific trade-offs between reproduction and defense. To investigate Brachypodium distachyon acclimation strategies to combinations of salinity, drought and heat, we applied a system biology approach, integrating physiological, metabolic and transcriptional analyses. We analyzed the trade-offs among functional and performance traits, and their effects on plant fitness. A combination of drought and heat resulted in escape strategy, while under a combination of salinity and heat, plants exhibited avoidance strategy. On the other hand, under combinations of salinity and drought, with or without heat stress, plant fitness (i.e. germination rate of subsequent generation) was severely impaired. These results indicate that under combined stresses, plants life-history strategies were shaped by the limits of phenotypic and metabolic plasticity and the trade-offs between traits, thereby giving raise to distinct acclimations. Our findings provide a mechanistic understanding of plant acclimations to combinations of abiotic stresses and shed light on the different life-history strategies that can contribute to grass fitness and possibly to their dispersion under changing environments.

plant biology

Mapping DNA sequence to transcription factor binding energy in vivo

Despite the central importance of transcriptional regulation in systems biology, it has proven difficult to determine the regulatory mechanisms of individual genes, let alone entire gene networks. It is particularly difficult to analyze a promoter sequence and identify the locations, regulatory roles, and energetic properties of binding sites for transcription factors and RNA polymerase. In this work, we present a strategy for interpreting transcriptional regulatory sequences using in vivo methods (i.e. the massively parallel reporter assay Sort-Seq) to formulate quantitative models that map a transcription factor binding sites DNA sequence to transcription factor-DNA binding energy. We use these models to predict the binding energies of transcription factor binding sites to within 1 kBT of their measured values. We further explore how such a sequence-energy mapping relates to the mechanisms of trancriptional regulation in various promoter contexts. Specifically, we show that our models can be used to design specific induction responses, analyze the effects of amino acid mutations on DNA sequence preference, and determine how regulatory context affects a transcription factors sequence specificity.

biophysics

Cellular heterogeneity in pressure and growth emerges from tissue topology and geometry

Cell-to-cell heterogeneity prevails in many biological systems, although its origin and function are often unclear. Cell hydrostatic pressure, alias turgor pressure, is essential in physiology and morphogenesis, and its spatial variations are often overlooked. Here, based on a mathematical model describing cell mechanics and water movement in a plant tissue, we predict that cell pressure anticorrelates with cell neighbour number. Using atomic force microscopy, we confirm this prediction in the Arabidopsis shoot apical meristem, a population of stem cells that generate all plant aerial organs. Pressure is predicted to correlate either positively or negatively with cellular growth rate depending on osmotic drive, cell wall extensibility, and hydraulic conductivity. The meristem exhibits one of these two regimes depending on conditions, suggesting that, in this tissue, water conductivity is non-negligible in growth control. Our results illustrate links between local topology, cell mechanical state and cell growth, with potential roles in tissue homeostasis.

biophysics

Pinned, locked, pushed, and pulled traveling waves in structured environments

Traveling fronts describe the transition between two alternative states in a great number of physical and biological systems. Examples include the spread of beneficial mutations, chemical reactions, and the invasions by foreign species. In homogeneous environments, the alternative states are separated by a smooth front moving at a constant velocity. This simple picture can break down in structured environments such as tissues, patchy landscapes, and microfluidic devices. Habitat fragmentation can pin the front at a particular location or lock invasion velocities into specific values. Locked velocities are not sensitive to moderate changes in dispersal or growth and are determined by the spatial and temporal periodicity of the environment. The synchronization with the environment results in discontinuous fronts that propagate as periodic pulses. We characterize the transition from continuous to locked invasions and show that it is controlled by positive density-dependence in dispersal or growth. We also demonstrate that velocity locking is robust to demographic and environmental fluctuations and examine stochastic dynamics and evolution in locked invasions.

ecology

Multiobjective Strain Design: A Framework for Modular Cell Engineering

Diversity of cellular metabolism can be harnessed to produce a large space of molecules. However, development of optimal strains with high product titers, rates, and yields required for industrial production is laborious and expensive. To accelerate the strain engineering process, we have recently introduced a modular cell design concept that enables rapid generation of optimal production strains by systematically assembling a modular cell with an exchangeable production module(s) to produce target molecules efficiently. In this study, we formulated the modular cell design concept as a general multiobjective optimization problem with flexible design objectives derived from mass action. We developed algorithms and an associated software package, named ModCell2 to implement the design. We demonstrated that ModCell2 can systematically identify genetic modifications to design modular cells that can couple with a variety of production modules and exhibit a minimal tradeoff among modularity, performance, and robustness. Analysis of the modular cell designs revealed both intuitive and complex metabolic architectures enabling modular production of these molecules. We envision ModCell2 provides a powerful tool to guide modular cell engineering and sheds light on modular design principles of biological systems.

bioengineering

Isolation, Development, and Genomic Analysis of Bacillus megaterium SR7 for Growth and Metabolite Production Under Supercritical Carbon Dioxide

Supercritical carbon dioxide (scCO2) is an attractive substitute for conventional organic solvents due to its unique transport and thermodynamic properties, its renewability and labile nature, and its high solubility for compounds such as alcohols, ketones and aldehydes. However, biological systems that use scCO2 are mainly limited to in vitro processes due to its strong inhibition of cell viability and growth. To solve this problem, we used a bioprospecting approach to isolate a microbial strain with the natural ability to grow while exposed to scCO2. Enrichment culture and serial passaging of deep subsurface fluids from the McElmo Dome scCO2 reservoir in aqueous media under scCO2 headspace enabled the isolation of spore-forming strain Bacillus megaterium SR7. Sequencing and analysis of the complete 5.51 Mbp genome and physiological characterization revealed the capacity for facultative anaerobic metabolism, including fermentative growth on a diverse range of organic substrates. Supplementation of growth medium with O_SCPCAPLC_SCPCAP-alanine for chemical induction of spore germination significantly improved growth frequencies and biomass accumulation under scCO2 headspace. Detection of endogenous fermentative compounds in cultures grown under scCO2 represents the first observation of bioproduct generation and accumulation under this condition. Culturing development and metabolic characterization of B. megaterium SR7 represent initial advancements in the effort towards enabling exploitation of scCO2 as a sustainable solvent for in vivo bioprocessing.

microbiology

Unsupervised phenotypic analysis of cellular images with multi-scale convolutional neural networks

Large-scale cellular imaging and phenotyping is a widely adopted strategy for understanding biological systems and chemical perturbations. Quantitative analysis of cellular images for identifying phenotypic changes is a key challenge within this strategy, and has recently seen promising progress with approaches based on deep neural networks. However, studies so far require either pre-segmented images as input or manual phenotype annotations for training, or both. To address these limitations, we have developed an unsupervised approach that exploits the inherent groupings within cellular imaging datasets to define surrogate classes that are used to train a multi-scale convolutional neural network. The trained network takes as input full-resolution microscopy images, and, without the need for segmentation, yields as output feature vectors that support phenotypic profiling. Benchmarked on two diverse benchmark datasets, the proposed approach yields accurate phenotypic predictions as well as compound potency estimates comparable to the state-of-the-art. More importantly, we show that the approach identifies novel cellular phenotypes not included in the manual annotation nor detected by previous studies.\n\nAuthor summaryCellular microscopy images provide detailed information about how cells respond to genetic or chemical treatments, and have been widely and successfully used in basic research and drug discovery. The recent breakthrough of deep learning methods for natural imaging recognition tasks has triggered the development and application of deep learning methods to cellular images to understand how cells change upon perturbation. Although successful, deep learning studies so far either can only take images of individual cells as input or require human experts to label a large amount of images. In this paper, we present an unsupervised deep learning approach that, without any human annotation, analyzes directly full-resolution microscopy images displaying typically hundreds of cells. We apply the approach to two benchmark datasets, and show that the approach identifies novel visual phenotypes not detected by previous studies.

bioinformatics

McImpute: Matrix completion based imputation for single cell RNA-seq data

MotivationSingle cell RNA sequencing has been proved to be revolutionary for its potential of zooming into complex biological systems. Genome wide expression analysis at single cell resolution, provides a window into dynamics of cellular phenotypes. This facilitates characterization of transcriptional heterogeneity in normal and diseased tissues under various conditions. It also sheds light on development or emergence of specific cell populations and phenotypes. However, owing to the paucity of input RNA, a typical single cell RNA sequencing data features a high number of dropout events where transcripts fail to get amplified.\n\nResultsWe introduce mcImpute, a low-rank matrix completion based technique to impute dropouts in single cell expression data. On a number of real datasets, application of mcImpute yields significant improvements in separation of true zeros from dropouts, cell-clustering, differential expression analysis, cell type separability, performance of dimensionality reduction techniques for cell visualization and gene distribution.\n\nAvailability and Implementationhttps://github.com/aanchalMongia/McImpute_scRNAseq

bioinformatics

Transcriptional regulatory mechanisms of fibrosis development in mouse lung tissue exposed to carbon nanotubes

BackgroundCarbon nanotubes (CNTs) usage has rapidly increased in the last few decades due to their unique properties, exploited in various industrial and commercial products. Certain types of CNTs cause adverse health effects, including chronic inflammation and fibrosis. Despite the large number of in vitro and in vivo studies evaluating these effects, many important questions remain unanswered due to a lack of mechanistic understanding of how CNTs induce cellular stress responses. In order to predict CNT toxicity, it is important to understand which transcriptional programs are specifically activated in response to CNTs, and what similarities and differences exist in relation to other toxic inducers exerting similar adverse effects.\n\nResultsA systems biology approach was applied to reveal complex interactions at the molecular level in mouse lung tissue in response to different fibrosis inducers: two types of multi-walled CNTs, NM-401 and NRCWE-26, and bleomycin (BLM). Based on mRNA gene expression profiles, we inferred gene regulatory networks (GRNs) to capture functional hierarchical regulatory structures between genes and their regulators. We found that activities of the transcription factors (TFs) Myc, Arid5a and Mxd1 were associated with the regulation of cytokine transcription in response to CNTs, while in response to BLM treatment, Myc was associated with p53 signaling. TF Litaf was identified as the essential regulator for noncanonical signaling of TLR2/4 driven by CNTs. Despite the different nature of the lung injury caused by CNTs and BLM, we identified common stress response modules, that included DNA damage (TFs: E2f8, E2f1, Foxm1), M1/M2 macrophage polarization (TF: Mafb), Interferon response (TFs: Irf7, Stat2 and Irf9) for all agents.\n\nConclusionsThese results suggest that the reconstruction and analysis of TF-centric gene interaction networks can reveal key targets and regulators of cellular stress responses to toxic agents.

pharmacology and toxicology

Spatio-Temporal Network Dynamics of Genes Underlying Schizophrenia

Schizophrenia (SZ) is a debilitating mental illness with multigenic etiology and high heritability. Despite extensive genetic studies the molecular etiology stays enigmatic. A systems biology study had suggested a protein-protein interaction (PPI) network for SZ with 504 novel PPIs amongst which several genes happen to be drug targets of existing FDA approved drugs. Although the PPI network presented all possible pairs of interactions (known and novel), it lacks a spatio-temporal information. The onset of psychiatric disorders is predominantly in adolescent and young adult stages, often accompanied by subtle structural abnormalities in multiple regions of the brain. Hence, there is a need to redefine the generic PPI network as a function of time (developmental stages) and space (brain regions). The availability of BrainSpan atlas data allowed us to redefine the SZ interactome as a function of space and time. The absence of non-synonymous variants in centenarians and non-psychiatric ExAC database allowed us to identify the variants of criticality. The expression of candidate genes in different brain regions and during developmental stages, responsible for cognitive processes as well as the onset of disease were studied. A subset of novel interactors detected in the network was further validated using gene-expression data of psychiatric postmortem brains. From the long list of drug targets proposed from the interactome study and based on the microarray gene-expression results, we have shortlisted a probable subset of 10 drug targets (targeted by 34 FDA approved drugs) coalescing into 81 biological pathways, that could be potentially repurposed for neuropsychiatric disorders.

bioinformatics

ROSeq: A rank based approach to modelling gene expression in single cells

1Systematic delineation of complex biological systems is an ever-challenging and resource-intensive process. Single cell transcriptomics allows us to study cell-to-cell variability in complex tissues at an unprecedented resolution. Accurate modeling of gene expression plays a critical role in the statistical determination of tissue-specific gene expression patterns. In the past few years, considerable efforts have been made to identify appropriate parametric models for single cell expression data. The zero-inflated version of Poisson/Negative Binomial and Log-Normal distributions have emerged as the most popular alternatives due to their ability to accommodate high dropout rates, as commonly observed in single cell data. While the majority of the parametric approaches directly model expression estimates, we explore the potential of modeling expression-ranks, as robust surrogates for transcript abundance. Here we examined the performance of the Discrete Generalized Beta Distribution (DGBD) on real data and devised a Wald-type test for comparing gene expression across two phenotypically divergent groups of single cells. We performed a comprehensive assessment of the proposed method, to understand its advantages as compared to some of the existing best practice approaches. Besides striking a reasonable balance between Type 1 and Type 2 errors, we concluded that ROSeq, the proposed differential expression test is exceptionally robust to expression noise and scales rapidly with increasing sample size. For wider dissemination and adoption of the method, we created an R package called ROSeq, and made it available on the Bioconductor platform.

genomics

Chronic inflammatory pain drives alcohol drinking in a sex-dependent manner

Sex differences in chronic pain and alcohol abuse are not well understood. The development of rodent models is imperative for investigating the underlying changes behind these pathological states. However, past attempts have failed to produce drinking outcomes similar to those reported in humans. In the present study, we investigated whether hind paw treatment with the inflammatory agent Complete Freunds Adjuvant (CFA) could generate hyperalgesia and alter alcohol consumption in male and female C57BL/6J mice. CFA treatment led to greater nociceptive sensitivity for both sexes in the Hargreaves test, and increased alcohol drinking for males in a continuous access two-bottle choice (CA2BC) paradigm. Regardless of treatment, female mice exhibited greater alcohol drinking than males. Following a 2-hour terminal drinking session, CFA treatment failed to produce changes in alcohol drinking, blood ethanol concentration (BEC), and plasma corticosterone (CORT) for both sexes. 2-hr alcohol consumption and CORT was higher in females than males, irrespective of CFA treatment. Taken together, these findings have established that male mice are more susceptible to escalations in alcohol drinking when undergoing pain, despite higher levels of total alcohol drinking and CORT in females. Furthermore, the exposure of CFA-treated C57BL/6J mice to the CA2BC drinking paradigm has proven to be a useful model for studying the relationship between chronic pain and alcohol abuse. Future applications of the CFA/CA2BC model should incorporate manipulations of stress signaling and other related biological systems to improve our mechanistic understanding of pain and alcohol interactions.

animal behavior and cognition

A Robust Method to Estimate the Largest Lyapunov Exponent of Noisy Signals: A Revision to the Rosenstein’s Algorithm

AimThis study proposed a revision to the Rosensteins method of numerical calculation of largest Lyapunov exponent (LyE) to make it more robust to noise.\n\nMethodsTo this aim, the effect of increasing number of initial neighboring points on the LyE value was investigated and compared to the values obtained by filtering the time series. Both simulated (Lorenz and passive dynamic walker) and experimental (human walking) time series were used to calculate LyE. The number of initial neighbors used to calculate LyE for all time series was 1 (the original Rosensteins method), 2, 3, 4, 5, 10, 15, 20, 25, and 30 data points.\n\nResultsThe results demonstrated that the LyE graph reached a plateau at the 15-point neighboring condition inferring that the LyE values calculated using at least 15 neighboring points were consistent and reliable.\n\nConclusionThe proposed method could be used to calculate LyE more reliably in experimental time series acquired from biological systems where noise is omnipresent.

biophysics

Genome-wide repressive capacity of promoter DNA methylation is revealed through epigenomic manipulation

The scientific community is increasingly embracing open science. This growing commitment to open science should be applauded and encouraged, especially when it occurs voluntarily and prior to peer review. Thanks to other researchers dedication to open science, we have had the privilege of conducting a reanalysis of a landmark experiment published as a preprint with data made available in a public repository. The study in question found that promoter DNA methylation is frequently insufficient to induce transcriptional repression, which appears to contradict a large body of observational studies showing a strong association between DNA methylation and gene expression. This study was the first to evaluate whether forcibly methylating thousands of DNA promoter regions is sufficient to suppress gene expression. The authors data analysis did not find a strong relationship between promoter methylation and transcriptional repression. However, their analyses did not make full use of statistical inference and applied a normalization technique that removes global differences that are representative of the actual biological system. Here we reanalyze the data with an approach that includes statistical inference of differentially methylated regions, as well as a normalization technique that accounts for global expression differences. We find that forced DNA methylation of thousands of promoters overwhelmingly represses gene expression. In addition, we show that complementary epigenetic marks of active transcription are reduced as a result of DNA methylation. Finally, by studying whether these associations are sensitive to the CG density of promoters, we find no substantial differences in the association between promoters with and without a CG island. The code needed to reproduce are analysis is included in the public GitHub repository github.com/kdkorthauer/repressivecapacity.

genomics

The genetics of the mood disorder spectrum: genome-wide association analyses of over 185,000 cases and 439,000 controls

BackgroundMood disorders (including major depressive disorder and bipolar disorder) affect 10-20% of the population. They range from brief, mild episodes to severe, incapacitating conditions that markedly impact lives. Despite their diagnostic distinction, multiple approaches have shown considerable sharing of risk factors across the mood disorders.\n\nMethodsTo clarify their shared molecular genetic basis, and to highlight disorder-specific associations, we meta-analysed data from the latest Psychiatric Genomics Consortium (PGC) genome-wide association studies of major depression (including data from 23andMe) and bipolar disorder, and an additional major depressive disorder cohort from UK Biobank (total: 185,285 cases, 439,741 controls; non-overlapping N = 609,424).\n\nResultsSeventy-three loci reached genome-wide significance in the meta-analysis, including 15 that are novel for mood disorders. More genome-wide significant loci from the PGC analysis of major depression than bipolar disorder reached genome-wide significance. Genetic correlations revealed that type 2 bipolar disorder correlates strongly with recurrent and single episode major depressive disorder. Systems biology analyses highlight both similarities and differences between the mood disorders, particularly in the mouse brain cell types implicated by the expression patterns of associated genes. The mood disorders also differ in their genetic correlation with educational attainment - positive in bipolar disorder but negative in major depressive disorder.\n\nConclusionsThe mood disorders share several genetic associations, and can be combined effectively to increase variant discovery. However, we demonstrate several differences between these disorders. Analysing subtypes of major depressive disorder and bipolar disorder provides evidence for a genetic mood disorders spectrum.

genomics

Modular assembly of polysaccharide-degrading microbial communities in the ocean

Many complex biological systems such as metabolic networks can be divided into functional and organizational subunits, called modules, which provide the flexibility to assemble novel multi-functional hierarchies by a mix and match of simpler components. Here we show that polysaccharide-degrading microbial communities in the ocean can also assemble in a modular fashion. Using synthetic particles made of a variety of polysaccharides commonly found in the ocean, we showed that the particle colonization dynamics of natural bacterioplankton assemblages can be understood as the aggregation of species modules of two main types: a first module type made of narrow niche-range primary degraders, whose dynamics are controlled by particle polysaccharide composition, and a second module type containing broad niche-range, substrate-independent taxa whose dynamics are controlled by interspecific interactions, in particular cross-feeding via organic acids, amino acids and other metabolic byproducts. As a consequence of this modular logic, communities can be predicted to assemble by a sum of substrate-specific primary degrader modules, one for each complex polysaccharide in the particle, connected to a single broad-niche range consumer module. We validate this model by showing that a linear combination of the communities on single-polysaccharide particles accurately predicts community composition on mixed-polysaccharide particles. Our results suggest thus that the assembly of heterotrophic communities that degrade complex organic materials follow simple design principles that can be exploited to engineer heterotrophic microbiomes.

microbiology