Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Systems Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Assessing the impact of high-throughput sequencing strategy and model complexity on ABC-inferred demographic history in mussels

Genome-scale diversity data are increasingly available in a variety of biological systems, and can be used to reconstruct the past evolutionary history of species divergence. However, extracting the full demographic information from these data is not trivial, and requires inferential methods that account for the diversity of coalescent histories throughout the genome. Here, we evaluate the potential and limitations of one such approach. We reexamine a well-known system of mussel sister species, using the joint site frequency spectrum (jSFS) of synonymous mutations computed either from exome capture or RNA-seq, in an Approximate Bayesian Computation (ABC) framework. We first assess the best sampling strategy (number of: individuals, loci, and bins in the jSFS), and show that model selection is robust to variation in the number of individuals and loci. In contrast, different binning choices when summarizing the joint site frequency spectrum, strongly affect the results: including classes of low and high frequency shared polymorphisms can more effectively reveal recent migration events. We then take advantage of the flexibility of ABC to compare more realistic models of speciation, including variation in migration rates through time (i.e. periodic connectivity) and across genes (i.e. genome-wide heterogeneity in migration rates). We show that these models were consistently selected as the most probable, suggesting that mussels have experienced a complex history of gene flow during divergence and that the species boundary is semi-permeable. Our work provides a comprehensive evaluation of ABC demographic inference in mussels based on the coding site frequency spectrum, and supplies guidelines for employing different sequencing techniques and sampling strategies. We emphasize, perhaps surprisingly, that inferences are less limited by the volume of data, than by the way in which they are analyzed.

evolutionary biology

Autoamplification and competition drive symmetry breaking: Initiation of centriole duplication by the PLK4-STIL network

Symmetry breaking, a central principle of physics, has been hailed as the driver of self-organization in biological systems in general and biogenesis of cellular organelles in particular, but the molecular mechanisms of symmetry breaking only begin to become understood. Centrioles, the structural cores of centrosomes and cilia, must duplicate every cell cycle to ensure their faithful inheritance through cellular divisions. Work in model organisms identified conserved proteins required for centriole duplication and found that altering their abundance affects centriole number. However, the biophysical principles that ensure that, under physiological conditions, only a single procentriole is produced on each mother centriole remain enigmatic. Here we propose a mechanistic biophysical model for the initiation of procentriole formation in mammalian cells. We posit that interactions between the master regulatory kinase PLK4 and its activator-substrate STIL form the basis of the procentriole initiation network. The model faithfully recapitulates the experimentally observed transition from PLK4 uniformly distributed around the mother centriole, the \"ring\", to a unique PLK4 focus, the \"spot\", that triggers the assembly of a new procentriole. This symmetry breaking requires a dual positive feedback based on autocatalytic activation of PLK4 and enhanced centriolar anchoring of PLK4-STIL complexes by phosphorylated STIL. We find that, contrary to previous proposals, in situ degradation of active PLK4 is insufficient to break symmetry. Instead, the model predicts that competition between transient PLK4 activity maxima for PLK4-STIL complexes explains both the instability of the PLK4 ring and formation of the unique PLK4 spot. In the model, strong competition at physiologically normal parameters robustly produces a single procentriole, while increasing overexpression of PLK4 and STIL weakens the competition and causes progressive addition of procentrioles in agreement with experimental observations.

cell biology

Universality and predictability in molecular quantitative genetics

Molecular traits, such as gene expression levels or protein binding affinities, are increasingly accessible to quantitative measurement by modern high-throughput techniques. Such traits measure molecular functions and, from an evolutionary point of view, are important as targets of natural selection. We review recent developments in evolutionary theory and experiments that are expected to become building blocks of a quantitative genetics of molecular traits. We focus on universal evolutionary characteristics: these are largely independent of a traits genetic basis, which is often at least partially unknown. We show that universal measurements can be used to infer selection on a quantitative trait, which determines its evolutionary mode of conservation or adaptation. Furthermore, universality is closely linked to predictability of trait evolution across lineages. We argue that universal trait statistics extends over a range of cellular scales and opens new avenues of quantitative evolutionary systems biology.

Evolutionary Biology

Estimating cellular pathways from an ensemble of heterogeneous data sources

Building better models of cellular pathways is one of the major challenges of systems biology and functional genomics. There is a need for methods to build on established expert knowledge and reconcile it with results of high-throughput studies. Moreover, the available data sources are heterogeneous and need to be combined in a way specific for the part of the pathway in which they are most informative. Here, we present a compartment specific strategy to integrate edge, node and path data for the refinement of a network hypothesis. Specifically, we use a local-move Gibbs sampler for refining pathway hypotheses from a compendium of heterogeneous data sources, including novel methodology for integrating protein attributes. We demonstrate the utility of this approach in a case study of the pheromone response MAPK pathway in the yeast S. cerevisiae.

Genomics

BoCluSt: bootstrap clustering stability algorithm for community detection in networks

The identification of modules or communities of related variables is a key step in the analysis and modelling of biological systems. Many module identification procedures are available, but few of these can determine the module partitions best fitting a given dataset in the absence of previous information, in an unsupervised way, and when the links between variables have different weights. Here I propose such a procedure, which uses the stability under bootstrap resampling of different alternative module structures as a criterion to identify the structure best fitting to a set of variables. In its present implementation, the procedure uses linear correlations as link weights. Computer simulations show that the procedure is useful for problems involving moderate numbers of variables, such as those commonly found in gene regulation cascades or metabolic pathways, and also that it can detect hierarchical network structures, in which modules are composed of smaller sub modules. The procedure becomes less practical as the number of variables increases, due to increases in processing time.The proposed procedure may be a valuable and robust network analysis tool. Because it is based on comparing the amount of evidence for different module partitions structures, this procedure may detect the existence of hierarchical network structures.

Bioinformatics

FORGE: multivariate calculation of gene-wide p-values from Genome-Wide Association Studies Authors and Affiliations

Genome-wide association studies (GWAS) have proven a valuable tool to explore the genetic basis of many traits. However, many GWAS lack statistical power and the commonly used single-point analysis method needs to be complemented to enhance power and interpretation. Multivariate region or gene-wide association are an alternative, allowing for identification of disease genes in a manner more robust to allelic heterogeneity. Gene-based association also facilitates systems biology analyses by generating a single p-value per gene. We have designed and implemented FORGE, a software suite which implements a range of methods for the combination of p-values for the individual genetic variants within a gene or genomic region. The software can be used with summary statistics (marker ids and p-values) and accepts as input the result file formats of commonly used genetic association software. When applied to a study of Crohns disease susceptibility, it identified all genes found by single SNP analysis and additional genes identified by large independent metaanalysis. FORGE p-values on gene-set analyses highlighted association with the Jak-STAT and cytokine signalling pathways, both previously associated with CD. We highlight the softwares main features, its future development directions and provide a comparison with alternative available software tools. FORGE can be freely accessed at https://github.com/inti/FORGE.

Genetics

Unsupervised data-driven stratification of mentalizing heterogeneity in autism

Individuals affected by autism spectrum conditions (ASC) are considerably heterogeneous. Novel approaches are needed to parse this heterogeneity to enhance precision in clinical and translational research. Applying a clustering approach taken from genomics and systems biology on two large independent cognitive datasets of adults with and without ASC (n=715; n=251), we find replicable evidence for 5 discrete ASC subgroups that are highly differentiated in item-level performance on an explicit mentalizing task tapping ability to read complex emotion and mental states from the eye region of the face (Reading the Mind in the Eyes Test; RMET). Three subgroups comprising 42-65% of ASC adults show evidence for large impairments (Cohens d = -1.03 to -11.21), while other subgroups are effectively unimpaired. These findings delineate robust natural subdivisions within the ASC population that may allow for more individualized inferences and accelerate research towards precision medicine goals.

Animal Behavior and Cognition

Towards Consensus Gene Ages

Correctly estimating the age of a gene or gene family is important for a variety of fields, including molecular evolution, comparative genomics, and phylogenetics, and increasingly for systems biology and disease genetics. However, most studies use only a point estimate of a genes age, neglecting the substantial uncertainty involved in this estimation. Here, we characterize this uncertainty by investigating the effect of algorithm choice on gene-age inference and calculate consensus gene ages with attendant error distributions for a variety of model eukaryotes. We use thirteen orthology inference algorithms to create gene-age datasets and then characterize the error around each age-call on a per-gene and per-algorithm basis. Systematic error was found to be a large factor in estimating gene age, suggesting that simple consensus algorithms are not enough to give a reliable point estimate. We also found that different sources of error can affect downstream analyses, such as gene ontology enrichment. Our consensus gene-age datasets, with associated error terms, are made fully available at so that researchers can propagate this uncertainty through their analyses (https://github.com/marcottelab/Gene-Ages).

Genomics

Paradoxical signaling regulates structural plasticity in dendritic spines

Transient spine enlargement (3-5 min timescale) is an important event associated with the structural plasticity of dendritic spines. Many of the molecular mechanisms associated with transient spine en{-}largement have been identified experimentally. Here, we use a systems biology approach to construct a mathematical model of biochemical signaling and actin-mediated transient spine expansion in response to calcium-influx due to NMDA receptor activation. We have identified that a key feature of this signaling network is the paradoxical signaling loop. Paradoxical components act bifunctionally in signaling net{-}works and their role is to control both the activation and inhibition of a desired response function (protein activity or spine volume). Using ordinary differential equation (ODE)-based modeling, we show that the dynamics of different regulators of transient spine expansion including CaMKII, RhoA, and Cdc42 and the spine volume can be described using paradoxical signaling loops. Our model is able to capture the experimentally observed dynamics of transient spine volume. Furthermore, we show that actin remod{-}eling events provide a robustness to spine volume dynamics. We also generate experimentally testable predictions about the role of different components and parameters of the network on spine dynamics.

Neuroscience

Diagnostic assessments of student thinking about stochastic processes.

A number of research studies indicate that students often have difficulties in understanding the presence and/or the implications of stochastic processes within biological systems. While critical to a wide range of phenomena, the presence and implications of stochastic processes are rarely explicitly considered in the course of formal instruction. To help instructors identify gaps in student understanding, we have designed and tested six open source activities covering a range of scenarios, from death rates to noise in gene expression, that can be employed, alone or in combination, as diagnostics to reveal student thinking as a prelude to the presentation of stochastic processes within a course or a curriculum.

Scientific Communication and Education

Exposing the diversity of multiple infection patterns

Natural populations often have to cope with genetically distinct parasites that can coexist, or not, within the same hosts. Theoretical models addressing the evolution of virulence have considered two within host infection outcomes, namely superinfection and coinfection. The field somehow became limited by this dichotomy that does not correspond to an empirical reality, as other infection patterns, namely sets of within-host infection outcomes, are possible. We indeed formally prove there are 114 different infection patterns for the sole recoverable chronic infections caused by horizontally-transmitted microparasites. We afterwards highlight eight infection patterns using an explicit modelling of within-host dynamics that captures a large range of ecological interactions, five of which have been neglected so far. To clarify the terminology related to multiple infections, we introduce terms describing these new relevant patterns and illustrate them with existing biological systems. This characterisation of infection patterns opens new perspectives for understanding the epidemiology and the evolution of parasites.

Epidemiology

Ouija: Incorporating prior knowledge in single-cell trajectory learning using Bayesian nonlinear factor analysis

Pseudotime estimation from single-cell gene expression allows the recovery of temporal information from otherwise static profiles of individual cells. This pseudotemporal information can be used to characterise transient events in temporally evolving biological systems. Conventional algorithms typically emphasise an unsupervised transcriptome-wide approach and use retrospective analysis to evaluate the behaviour of individual genes. Here we introduce an orthogonal approach termed \"Ouija\" that learns pseudotimes from a small set of marker genes that might ordinarily be used to retrospectively confirm the accuracy of unsupervised pseudotime algorithms. Crucially, we model these genes in terms of switch-like or transient behaviour along the trajectory, allowing us to understand why the pseudotimes have been inferred and learn informative parameters about the behaviour of each gene. Since each gene is associated with a switch or peak time the genes are effectively ordered along with the cells, allowing each part of the trajectory to be understood in terms of the behaviour of certain genes. In the following we introduce our model and demonstrate that in many instances a small panel of marker genes can recover pseudotimes that are consistent with those obtained using the entire transcriptome. Furthermore, we show that our method can detect differences in the regulation timings between two genes and identify \"metastable\" states - discrete cell types along the continuous trajectories - that recapitulate known cell types. Ouija therefore provides a powerful complimentary approach to existing whole transcriptome based pseudotime estimation methods. An open source implementation is available at http://www.github.com/kieranrcampbell/ouija as an R package and at http://www.github.com/kieranrcampbell/ouijaflow as a Python/TensorFlow package.

Bioinformatics

DIABLO - an integrative, multi-omics, multivariate method for multi-group classification

Systems biology approaches, leveraging multi-omics measurements, are needed to capture the complexity of biological networks while identifying the key molecular drivers of disease mechanisms. We present DIABLO, a novel integrative method to identify multi-omics biomarker panels that can discriminate between multiple phenotypic groups. In the multi-omics analyses of simulated and real-world datasets, DIABLO resulted in superior biological enrichment compared to other integrative methods, and achieved comparable predictive performance with existing multi-step classification schemes. DIABLO is a versatile approach that will benefit a diverse range of research areas, where multiple high dimensional datasets are available for the same set of specimens. DIABLO is implemented along with tools for model selection, and validation, as well as graphical outputs to assist in the interpretation of these integrative analyses (http://mixomics.org/).

Bioinformatics

One-cell Doubling Evaluation by Living Arrays of Yeast, ODELAY!

Cell growth is a complex phenotype widely used in systems biology to gauge the impact of genetic and environmental perturbations. Due to the magnitude of genome-wide studies, resolution is often sacrificed in favor of throughput, creating a demand for scalable, time-resolved, quantitative methods of growth assessment. We present ODELAY (One-cell Doubling Evaluation by Living Arrays of Yeast), an automated and scalable growth analysis platform. High measurement density and single cell resolution provide a powerful tool for large-scale multiparameter growth analysis based on the modeling of microcolony expansion on solid media. Pioneered in yeast but applicable to other colony forming organisms, ODELAY extracts the three key growth parameters (lag time, doubling time, and carrying capacity) that define microcolony expansion from single cells, simultaneously permitting the assessment of population heterogeneity. The utility of ODELAY is illustrated using yeast mutants, revealing a spectrum of phenotypes arising from single and combinatorial growth parameter perturbations.

Microbiology

CH···O Interactions Are Not the Cause of Trends in Reactivity and Secondary Kinetic Isotope Effects for Enzymatic SN2 Methyl Transfer Reactions

Compaction mattersSN2 substitution represents an important class of reaction for both chemical and biological systems. The ability to assess enzymatic transition state structure within this class of reaction remains a major experimental challenge. Here, we comment on and compare the relative impact of compaction along the axis of reaction to the impact of an orthogonal CH{middle dot}{middle dot}{middle dot}O hydrogen bonding interaction. The latter is concluded to play a limited role in determining relative reaction rates and secondary KIEs derived from experimental structure-activity correlations.

Biochemistry

Reductive Analytics on Big MS Data leads to tremendous reduction in time for peptide deduction

In this paper we present a feasibility of using a data-reductive strategy for analyzing big MS data. The proposed method utilizes our reduction algorithm MS-REDUCE and peptide deduction is accomplished using Tide with hiXcorr. Using this approach we were able to process 1 million spectra in under 3 hours. Our results showed that running peptide deduction with smaller amount of selected peaks made the computations much faster and scalable with increasing resolution of MS data. Quality assessment experiments performed on experimentally generated datasets showed good quality peptide matches can be made using the reduced datasets. We anticipate that the proteomics and systems biology community will widely adopt our reductive strategy due to its efficacy and reduced time for analysis.

Bioinformatics

Purifying selection provides buffering of the natural variation co-expression network in a forest tree species

Several studies have investigated general properties of the genetic architecture of gene expression variation. Most of these used controlled crosses and it is unclear whether their findings extend to natural populations. Furthermore, systems biology has established that biological networks are buffered against large effect mutations, but there remains little data resolving this with natural variation of gene expression. Here we utilise RNA-Sequencing to assay gene expression in winter buds undergoing bud flush in a natural population of Populus tremula. We performed expression Quantitative Trait Locus (eQTL) mapping and identified 164,290 significant eQTLs associating 6,241 unique genes (eGenes) with 147,419 unique SNPs (eSNPs). We found approximately four times as many local as distant eQTLs, which had significantly higher effect size. eQTLs were primarily located in regulatory regions of genes (UTRs or flanking regions), regardless of whether they were local or distant. We used the gene expression data to infer a co-expression network and investigated to what degree eQTLs could explain the structure of the network: eGenes were present in the core of 28 of 38 network modules, however, eGenes were overall underrepresented in cores and overrepresented in the periphery of the network, with a negative correlation between eQTL effect size and network connectivity. We also observed a negative correlation between eQTL effect size and allele frequency and found that core genes have experienced stronger selective constraint. Our integrated genetics and genomics results suggest that prevalent purifying selection is the primary mechanism underlying the genetic architecture of natural variation in gene expression in P. tremula and that highly connected network hubs are buffered against deleterious effects as a result of regulation by numerous eSNPs, each of minor effect.\n\nAuthor summaryNumerous studies have shown that many genomic polymorphisms contributing to phenotypic variation are located outside of protein coding regions, suggesting that they act by modulating gene expression. Furthermore, phenotypes are seldom explained by individual genes, but rather emerge from networks of interacting genes. The effect of regulatory variants and the interaction of genes can be described by co-expression networks, which are known to contain a small number of highly connected nodes and many more lowly connected nodes, making them robust to random mutation. While previous studies have examined the genetic architecture of gene expression variation, few were performed in natural populations with fewer still integrating the co-expression network.\n\nWe undertook a study using a natural population of European aspen (Populus tremula), showing that expression variance is substantially smaller among individuals than between tissues within the same individual, suggesting that stabilizing selection may act to restrict the scale of expression variation. We further show that highly connected genes within the co-expression network are associated with polymorphisms of lower than average effect size, suggesting purifying selection. These genes are therefore buffered against large expression modulation, providing a mechanistic explanation of how network robustness is created and maintained at the population level.

Genetics

Identifying cis-mediators for trans-eQTLs across many human tissues using genomic mediation analysis

The impact of inherited genetic variation on gene expression in humans is well-established. The majority of known expression quantitative trait loci (eQTLs) impact expression of local genes (cis-eQTLs). More research is needed to identify effects of genetic variation on distant genes (trans-eQTLs) and understand their biological mechanisms. One common trans-eQTLs mechanism is \"mediation\" by a local (cis) transcript. Thus, mediation analysis can be applied to genome-wide SNP and expression data in order to identify transcripts that are \"cis-mediators\" of trans-eQTLs, including those \"cis-hubs\" involved in regulation of many trans-genes. Identifying such mediators helps us understand regulatory networks and suggests biological mechanisms underlying trans-eQTLs, both of which are relevant for understanding susceptibility to complex diseases. The multi-tissue expression data from the Genotype-Tissue Expression (GTEx) program provides a unique opportunity to study cis-mediation across human tissue types. However, the presence of complex hidden confounding effects in biological systems can make mediation analyses challenging and prone to confounding bias, particularly when conducted among diverse samples. To address this problem, we propose a new method: Genomic Mediation analysis with Adaptive Confounding adjustment (GMAC). It enables the search of a very large pool of variables, and adaptively selects potential confounding variables for each mediation test. Analyses of simulated data and GTEx data demonstrate that the adaptive selection of confounders by GMAC improves the power and precision of mediation analysis. Application of GMAC to GTEx data provides new insights into the observed patterns of cis-hubs and trans-eQTL regulation across tissue types.

Genomics