Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,657 records · Page 92Linked to original sources

Antioxidant activity and phycoremediation ability of four cyanobacterial isolates obtained from a stressed aquatic system

Cyanobacteria are natural enormous sources of various biologically active compounds with great contributions in different industries. This study aimed to introduce molecular and biochemical characterization for four novel cyanobacterial isolates obtained from Egyptian wastewater canals. Besides, In vitro biological activity of these isolates and their potential ability to take up nutrients and heavy metals from wastewater were examined. The obtained accession numbers were KY250420.1, KY321359.1, KY296359.1 and KU373076.1 for Nostoc calcicola, Leptolyngbya sp, Nostoc sp, and Nostoc sp, respectively. The isolate Leptolyngbya sp (KY321359.1) showed the lowest identity (90%) with other deposited sequences in database. While the isolate Nostoc sp (KU373076.1) showed the highest total phenolic content as well as the highest levels of caffeic, ferulic and gallic acids. Consequently, it appeared the highest antioxidant scavenging activity. All cyanobacterial isolates revealed potent ability to take up nutrients and heavy metals from wastewater. Generally, this study provides a taxonomic and molecular evidence for four novel cyanobacterial isolates with antioxidant activity and potent phycoremediation ability.

genetics

Regime shifts, alternative states and hysteresis in the Sarracenia microecosystem

Changes in environmental conditions can lead to rapid shifts in the state of an ecosystem (\"regime shifts\"), which, even after the environment has returned to previous conditions, subsequently recovers slowly to the previous state (\"hysteresis\"). Large spatial and temporal scales of dynamics, and the lack of frameworks linking observations to models, are challenges to understanding and predicting ecosystem responses to perturbations. The naturally-occurring microecosystem inside leaves of the northern pitcher plant (Sarracenia purpurea) exhibits oligotrophic and eutrophic states that can be induced by adding insect prey. Here, we further develop a model for simulating these dynamics, parameterize it using data from a prey addition experiment and conduct a sensitivity analysis to identify critical zones within the parameter space. Simulations illustrate that the microecosystem model displays regime shifts and hysteresis. Parallel results were observed in the plant itself after experimental enrichment with prey. Decomposition rate of prey was the main driver of system dynamics, including the time the system remains in an anoxic state and the rate of return to an oxygenated state. Biological oxygen demand in fluenced the shape of the systems return trajectory. The combination of simulated results, sensitivity analysis and use of empirical results to parameterize the model more precisely demonstrates that the Sarracenia microecosystem model displays behaviors qualitatively similar to models of larger ecological systems.

ecology

Double-digest RAD-sequencing: do wet and dry protocol parameters impact biological results?

O_LINext-generation sequencing technologies have opened a new era of research in genomics. Among these, restriction enzyme-based techniques such as restriction-site associated DNA sequencing (RADseq) or double-digest RAD-sequencing (ddRADseq) are now widely used in many population genomics fields. From DNA sampling to SNP calling, both wet and dry protocols have been discussed in the literature to identify key parameters for an optimal loci reconstruction.\nC_LIO_LIThe impact of these parameters on downstream analyses and biological results drawn from RADseq or ddRADseq data has however not been fully explored yet. In this study, we tackled this issue by investigating the effects of ddRADseq laboratory (i.e. wet protocol) and bioinformatics (i.e. dry protocol) settings on loci reconstruction and inferred biological signal at two evolutionary scale using two systems: a complex of butterfly species (Coenonympha sp.) and populations of Common beech (Fagus sylvatica).\nC_LIO_LIResults suggest an impact of wet protocol parameters (DNA quantity, number of PCR cycles during library preparation) on the number of recovered reads and SNPs, the number of unique alleles and individual heterozygosity. We also found that bioinformatic settings (i.e. clustering and minimum coverage thresholds) impact loci reconstruction (e.g. number of loci, mean coverage) and SNP calling (e.g. number of SNPs, heterozygosity). We however do not detect an impact of parameter settings on three types of analysis performed with ddRADseq data: measure of genetic differentiation, estimation of individual admixture, and demographic inferences. In addition, our work demonstrates the high reproducibility and low rate of genotyping inconsistencies of the ddRADseq protocol.\nC_LIO_LIThus, our study highlights the impact of wet parameters on ddRADseq protocol with strong consequences on experimental success and biological conclusions. Dry parameters affects loci reconstruction and descriptive statistics but not biological conclusion for the two studied systems. Overall, this study illustrates, with others, the relevance of ddRADseq for population and evolutionary genomics at the inter- or intraspecific scales.\nC_LI

molecular biology

Phosphorylation energy and nonlinear kinetics as key determinants for G2/M transition in fission yeast cell cycle

The living cell is an open nonequilibrium biochemical system, where ATP hydrolysis serves as the energy source for a wide range of intracellular processes including the assurance for decision-making. In the fission yeast cell cycle, the transition from G2 phase to M phase is triggered by the activation of Cdc13/Cdc2 and Cdc25, and the deactivation of Wee1. Each of these three events involves a phosphorylation-dephosphorylation (PdP) cycle, and together they form a regulatory circuit with feedback loops. Almost all quantitative models for cellular networks in the past have invalid thermodynamics due to the assumption of irreversible enzyme kinetics. We constructed a thermodynamically realistic kinetic model of the G2/M circuit, and show that the phosphorylation energy ({Delta}G), which is determined by the cellular ATP/ADP ratio, critically controls the dynamics and the bistable nature of Cdc2 activation. Using fission yeast nucleoplasmic extract (YNPE), we are able to experimentally verify our model prediction that increased {Delta}G, being synergistic to the accumulation of Cdc13, drives the activation of Cdc2. Furthermore, Cdc2 activation exhibits bistability and hysteresis in response to changes in phosphorylation energy. These findings suggest that adequate maintenance of phosphorylation energy ensures the bistability and robustness of the activation of Cdc2 in the G2/M transition. Free energy might play a widespread role in biological decision-making processes, connecting thermodynamics with information processing in biology.

systems biology

System-wide automatic extraction of functional signatures in Pseudomonas aeruginosa with eADAGE

Cross experiment comparisons in public data compendia are challenged by unmatched conditions and technical noise. The ADAGE method, which performs unsupervised integration with neural networks, can effectively identify biological patterns, but because ADAGE models, like many neural networks, are over-parameterized, different ADAGE models perform equally well. To enhance model robustness and better build signatures consistent with biological pathways, we developed an ensemble ADAGE (eADAGE) that integrated stable signatures across models. We applied eADAGE to a Pseudomonas aeruginosa compendium containing experiments performed in 78 media. eADAGE revealed a phosphate starvation response controlled by PhoB. While we expected PhoB activity in limiting phosphate conditions, our analyses found PhoB activity in other media with moderate phosphate and predicted that a second stimulus provided by the sensor kinase, KinB, is required for PhoB activation in this setting. We validated this relationship using both targeted and unbiased genetic approaches. eADAGE, which captures stable biological patterns, enables cross-experiment comparisons that can highlight measured but undiscovered relationships.

Systems Biology

Hybrid systems approach to modeling stochastic dynamics of cell size

A ubiquitous feature of all living cells is their growth over time followed by division into two daughter cells. How a population of genetically identical cells maintains size homeostasis, i.e., a narrow distribution of cell size, is an intriguing fundamental problem. We model size using a stochastic hybrid system, where a cell grows exponentially over time and probabilistic division events are triggered at discrete time intervals. Moreover, whenever these events occur, size is randomly partitioned among daughter cells. We first consider a scenario, where a timer (i.e., cell-cycle clock) that measures the time since the last division event regulates cellular growth and the rate of cell division. Analysis reveals that such a timer-driven system cannot achieve size homeostasis, in the sense that, the cell-to-cell size variation grows unboundedly with time. To explore biologically meaningful mechanisms for controlling size we consider three different classes of models: i) a size-dependent growth rate and timer-dependent division rate; ii) a constant growth rate and size-dependent division rate and iii) a constant growth rate and division rate that depends both on the cell size and timer. We show that each of these strategies can potentially achieve bounded intercellular size variation, and derive closed-form expressions for this variation in terms of underlying model parameters. Finally, we discuss how different organisms have adopted the above strategies for maintaining cell size homeostasis.

Systems Biology

Identifying (un)controllable dynamical behavior with applications to biomolecular networks

We present a technique applicable in any dynamical framework to identify control-robust subsets of an interacting system. These robust subsystems, which we call stable modules, are characterized by constraints on the variables that make up the subsystem. They are robust in the sense that if the defining constraints are satisfied at a given time, they remain satisfied for all later times, regardless of what happens in the rest of the system, and can only be broken if the constrained variables are externally manipulated. We identify stable modules as graph structures in an expanded network, which represents causal links between variable constraints. A stable module represents a system \"decision point\", or trap subspace. Using the expanded network, small stable modules can be composed sequentially to form larger stable modules that describe dynamics on the system level. Collections of large, mutually exclusive stable modules describe the systems repertoire of long-term behaviors. We implement this technique in a broad class of dynamical systems and illustrate its practical utility via examples and algorithmic analysis of two published biological network models. In the segment polarity gene network of Drosophila melanogaster, we obtain a state-space visualization that reproduces by novel means the four possible cell fates and predicts the outcome of cell transplant experiments. In the T-cell signaling network, we identify six signaling elements that determine the high-signal response and show that control of an element connected to them cannot disrupt this response.\n\nAuthor summaryWe show how to uncover the causal relationships between qualitative statements about the values of variables in ODE systems. We then show how these relationships can be used to identify subsystem behaviors that are robust to outside interventions. This informs potential system control strategies (e.g., in identifying drug targets). Typical analytical properties of biomolecular systems render them particularly amenable to our techniques. Furthermore, due to their often high dimension and large uncertainties, our results are particularly useful in biomolecular systems. We apply our methods to two quantitative biological models: the segment polarity gene network of Drosophila melanogaster and the T-cell signal transduction network.

systems biology

High-throughput cancer hypothesis testing with an integrated PhysiCell-EMEWS workflow

BackgroundCancer is a complex, multiscale dynamical system, with interactions between tumor cells and non-cancerous host systems. Therapies act on this combined cancer-host system, sometimes with unexpected results. Systematic investigation of mechanistic computational models can augment traditional laboratory and clinical studies, helping identify the factors driving a treatments success or failure. However, given the uncertainties regarding the underlying biology, these multiscale computational models can take many potential forms, in addition to encompassing high-dimensional parameter spaces. Therefore, the exploration of these models is computationally challenging. We propose that integrating two existing technologies--one to aid the construction of multiscale agent-based models, the other developed to enhance model exploration and optimization--can provide a computational means for high-throughput hypothesis testing, and eventually, optimization.\n\nResultsIn this paper, we introduce a high throughput computing (HTC) framework that integrates a mechanistic 3-D multicellular simulator (PhysiCell) with an extreme-scale model exploration platform (EMEWS) to investigate high-dimensional parameter spaces. We show early results in applying PhysiCell-EMEWS to 3-D cancer immunotherapy and show insights on therapeutic failure. We describe a generalized PhysiCell-EMEWS workflow for high-throughput cancer hypothesis testing, where hundreds or thousands of mechanistic simulations are compared against data-driven error metrics to perform hypothesis optimization.\n\nConclusionsWhile key notational and computational challenges remain, mechanistic agent-based models and high-throughput model exploration environments can be combined to systematically and rapidly explore key problems in cancer. These high-throughput computational experiments can improve our understanding of the underlying biology, drive future experiments, and ultimately inform clinical practice.

systems biology

Life Inside A Dinosaur Bone: A Thriving Microbiome

Fossils were long thought to lack original organic material, but the discovery of organic molecules in fossils and sub-fossils, thousands to millions of years old, has demonstrated the potential of fossil organics to provide radical new insights into the fossil record. How long different organics can persist remains unclear, however. Non-avian dinosaur bone has been hypothesised to preserve endogenous organics including collagen, osteocytes, and blood vessels, but proteins and labile lipids are unstable during diagenesis or over long periods of time. Furthermore, bone is porous and an open system, allowing microbial and organic flux. Some of these organics within fossil bone have therefore been identified as either contamination or microbial biofilm, rather than original organics. Here, we use biological and chemical analyses of Late Cretaceous dinosaur bones and sediment matrix to show that dinosaur bone hosts a diverse microbiome. Fossils and matrix were freshly-excavated, aseptically-acquired, and then analysed using microscopy, spectroscopy, chromatography, spectrometry, DNA extraction, and 16S rRNA amplicon sequencing. The fossil organics differ from modern bone collagen chemically and structurally. A key finding is that 16S rRNA amplicon sequencing reveals that the subterranean fossil bones host a unique, living microbiome distinct from that of the surrounding sediment. Even in the subsurface, dinosaur bone is biologically active and behaves as an open system, attracting microbes that might alter original organics or complicate the identification of original organics. These results suggest caution regarding claims of dinosaur bone soft tissue preservation and illustrate a potential role for microbial communities in post-burial taphonomy.

paleontology

Principles of studying a cell - a non-boastful paper for all molecular biologists

Studies of a cell rely on either observational approaches or perturbational/genetic approaches to define the contribution of a gene to specific cellular traits. It is unclear, however, under what circumstances each of the two approaches can be most successful and when they are doomed to fail. By analyzing over 500 complex traits of the yeast Saccharomyces cerevisiae we show that the trait relatedness to fitness determines the performance of observational approaches. Specifically, in traits subject to strong natural selection, genes identified using observational approaches are often highly coordinated in expression, such that the gene-trait associations are readily recognizable; in sharp contrast, the lack of such coordination in traits subject to weak selection leads to no detectable activity-trait associations for any individual genes and thus the failure of observational approaches. We further show that genetic approaches can be successful when the genes responsible for coordinating the target genes of observational approaches are perturbed. However, because the system-level cellular responses to a random mutation affect more or less every gene and consequently every trait, most genetic effects convey no trait-specific functional information for understanding the traits, which is particularly true for traits subject to weak selection.\n\nSignificance statementCell research is nearly exclusively based on empirical data obtained through either observational approaches or perturbational/genetic approaches. It is, however, increasingly clear that an analytical framework able to guide the empirical strategies is necessary to drive the field further ahead. This study analyzes ~500 complex traits of the yeast Saccharomyces cerevisiae and reveals the organizing principles of a cell. Specifically, a cell can be viewed as a factory, with each trait being the product of a production line operated directly by workers who are supervised by managers. For a cellular trait produced by many workers, the coordination level of the workers determines the performance of observational approaches. Meanwhile, the coordination of workers is realized by managers that are recruited and/or maintained by natural selection. Thus, observational approaches are expected to fail for traits subject to little selection, and genetic approaches can be successful only when the managers of fitness-tightly-coupled traits are perturbed. The manager-worker architecture built by natural selection explains well the origins of global epistasis and ubiquitous genetic effects, two major issues confusing current genetics and molecular and cellular biology, providing a clear guideline on how to study a cell.

Systems Biology

An integrative systems medicine approach to delineate complex genotype-phenotype associations in Autism Spectrum Disorder

The complex genetic architecture of Autism Spectrum Disorder (ASD) and its heterogeneous phenotype make molecular diagnosis and patient prognosis challenging tasks. To establish more precise genotype-phenotype correlations in ASD, we developed a novel machine learning integrative approach, which seeks to delineate associations between patients clinical profiles and disrupted biological processes inferred from their Copy Number Variants (CNVs) that span brain genes. Clustering analysis of relevant clinical measures from 2446 ASD cases in the Autism Genome Project identified two distinct phenotypic subgroups. Patients in these clusters differed significantly in ADOS-defined severity, adaptive behaviour profiles, intellectual ability and verbal status, the latter contributing the most for cluster stability and cohesion. Functional enrichment analysis of brain genes disrupted by CNVs in these ASD cases identified 15 statistically significant biological processes, including cell adhesion, neural development, cognition and polyubiquitination, in line with previous ASD findings. A Naive Bayes classifier, generated to predict the ASD phenotypic clusters from disrupted biological processes, achieved predictions with a high Precision (0.82) but low recall (0.39), for a subset of patients with higher biological Information Content scores. This study shows that milder and more severe clinical presentations can have distinct underlying biological mechanisms. It further highlights how machine learning approaches can reduce clinical heterogeneity using multidimensional clinical measures, and establish genotype-phenotype correlations in ASD. However, predictions are strongly dependent on patients information content. Findings are therefore a first step towards the translation of genetic information into clinically useful applications, but emphasize the need for larger datasets with very complete clinical and biological information.

systems biology

A longitudinal model of human neuronal differentiation for functional investigation of schizophrenia disease susceptibility

There is a pressing need for in vitro experimental systems that allow for interrogation of polygenic psychiatric disease risk to study the underlying biological mechanisms. We developed an analytical framework that integrates genome-wide disease risk from GWAS with longitudinal in vitro gene expression profiles of human neuronal differentiation. We demonstrate that the cumulative impact of risk loci of specific psychiatric disorders is significantly associated with genes that are differentially expressed across neuronal differentiation. We find significant evidence for schizophrenia, which is driven by a longitudinal synaptic gene cluster that is upregulated during differentiation. Our findings reveal that in vitro neuronal differentiation can be used to translate the polygenic architecture of schizophrenia to biologically relevant pathways that can be modeled in an experimental system. Overall, this work emphasizes the use of longitudinal in vitro transcriptomic signatures as a cellular readout and the application to the genetics of complex traits.

genetics

Beyond benchmark accuracy: machine-learning turnover-number predictors require system-level validation

Enzyme turnover numbers (kcat) are essential for kinetic models and enzyme-constrained genome-scale metabolic models (ecGEMs), but measured values are sparse and therefore increasingly estimated using machine learning (ML). Although these predictors are commonly evaluated by global regression metrics, their practical utility depends on how errors propagate through downstream models. We benchmarked six current kcat predictors on a curated BRENDA-derived dataset and five of them on EnzyExtract. To assess the influence of training-set proximity, we compared each benchmark dataset with the available training data for each predictor. We then used the predicted kcat values to parameterize ecGEMs of Saccharomyces cerevisiae and evaluated growth predictions across 19 conditions. We find that benchmark accuracy is moderate even on the BRENDA-derived dataset and drops sharply on EnzyExtract, where all predictors achieve R2 values of 0.20 or lower. This decline is accompanied by substantially lower overlap between the benchmark and training datasets, with exact sequence matches ranging from 24% to 78% for BRENDA, compared with 9% to 26% for EnzyExtract. However, that overlap alone does not explain differences in generalization across predictors. Moreover, downstream performance is also not explained by benchmark ranking. Across 19 conditions, none of the tool-specific ecGEMs consistently reproduces the experimentally observed variation in growth. In glucose minimal medium, the weakest benchmark performer yields the most accurate growth prediction in the downstream ecGEMs, whereas higher-ranked predictors produce larger deviations in growth. We trace this mismatch to localized errors at high-leverage positions in yeast's metabolic network, where underpredicted mitochondrial ADP/ATP carrier turnover numbers restrict adenine nucleotide exchange and impose an apparent limitation on cytosolic ATP supply. Relaxing this constraint shifts predicted growth toward the experimental reference. Thus, ML-derived kcat values can affect not only quantitative growth predictions but also the phenotype a mechanistic model appears to identify. These results argue for application-driven validation of biological parameter predictors in the downstream systems they are intended to support.

bioinformatics

The evolutionary advantage of heritable phenotypic heterogeneity

Phenotypic plasticity is an evolutionary driving force in diverse biological processes, including the adaptive immune system, the development of neoplasms, and the bacterial acquisition of drug resistance. It is essential, therefore, to understand the evolutionary advantage of an allele that confers cells the ability to express a range of phenotypes. Of particular importance is to understand how this advantage of phenotypic plasticity depends on the degree of heritability of non-genetically encoded phenotypes between generations, which can induce irreversible evolutionary changes in the population. Here, we study the fate of a new mutation that allows the expression of multiple phenotypic states, introduced into a finite population otherwise composed of individuals who can express only a single phenotype. We analyze the fixation probability of such an allele as a function of the strength of inter-generational phenotypic heritability, called memory, the variance of expressible phenotypes, the rate of environmental changes, and the population size. We find that the fate of a phenotypically plastic allele depends fundamentally on the environmental regime. In a constant environment, the fixation probability of a plastic allele always increases with the degree of phenotypic memory. In periodically fluctuating environments, by contrast, there is an optimum phenotypic memory that maximizes the probability of the plastic alleles fixation. This same optimum value of phenotypic memory also maximizes geometric mean fitness, in steady state. We interpret these results in the context of previous studies in an infinite-population framework. We also discuss the implications of our results for the design of therapies that can overcome resistance, in a variety of diseases.

Evolutionary Biology

Aberration-corrected high-NA open-top selective-plane illumination microscopy for biological imaging

AbstractSelective-plane illumination microscopy (SPIM) provides unparalleled advantages for volumetric imaging of living organisms over extended times. However, the spatial configuration of a SPIM system often limits its compatibility with many widely used biological sample holders such as multi-well chambers and plates. To solve this problem, we developed a high numerical aperture (NA) open-top configuration that places both the excitation and detection objectives on the opposite of the sample coverglass. We carried out a theoretical calculation to analyze the structure of the system-induced aberrations. We then experimentally compensated the system aberrations using adaptive optics combined with static optical components, demonstrating near-diffraction-limited performance in imaging fluorescently labeled cells.\n\n(C) 2017 Optical Society of America\n\nOCIS codes: (080.080) Geometric Optics; (110.0110) Imaging systems; (110.0180) Microscopy.

bioengineering

Regularization Improves the Robustness of Learned Sequence-to-Expression Models

Understanding of the gene regulatory activity of enhancers is a major problem in regulatory biology. The nascent field of sequence-to-expression modelling seeks to create quantitative models of gene expression based on regulatory DNA (cis) and cellular environmental (trans) contexts. All quantitative models are defined partially by numerical parameters, and it is common to fit these parameters to data provided by existing experimental results. However, the relative paucity of experimental data appropriate for this task, and lacunae in our knowledge of all components of the systems, results in problems often being under-specified, which in turn may lead to a situation where wildly different model parameterizations perform similarly well on training data. It may also lead to models being fit to the idiosyncrasies of the training data, without representing the more general process (overfitting).\n\nIn other contexts where parameter-fitting is performed, it is common to apply regularization to reduce overfitting. We systematically evaluated the efficacy of three types of regularization in improving the generalizability of trained sequence-to-expression models. The evaluation was performed in two types of cross-validation experiments: one training on D. melanogaster data and predicting on orthologous enhancers from related species, and the other cross-validating between four D. melanogaster neurogenic ectoderm enhancers, which are thought to be under control of the same transcription factors. We show that training with a combination of noise-injection, L1, and L2 regularization can drastically reduce overfitting and improve the generalizability of learned sequence-to-expression models. These results suggest that it may be possible to mitigate the tendency of sequence-to-expression models to overfit available data, thus improving predictive power and potentially resulting in models that provide better insight into underlying biological processes.

systems biology

Bayesian Energy Landscape Tilting: Towards Concordant Models of Molecular Ensembles

Predicting biological structure has remained challenging for systems such as disordered proteins that take on myriad conformations. Hybrid simulation/experiment strategies have been undermined by difficulties in evaluating errors from computa- tional model inaccuracies and data uncertainties. Building on recent proposals from maximum entropy theory and nonequilibrium thermodynamics, we address these issues through a Bayesian Energy Landscape Tilting (BELT) scheme for computing Bayesian \"hyperensembles\" over conformational ensembles. BELT uses Markov chain Monte Carlo to directly sample maximum-entropy conformational ensembles consistent with a set of input experimental observables. To test this framework, we apply BELT to model trialanine, starting from disagreeing simulations with the force fields ff96, ff99, ff99sbnmr-ildn, CHARMM27, and OPLS-AA. BELT incorporation of limited chemical shift and 3J measurements gives convergent values of the peptides , {beta}, and PPII conformational populations in all cases. As a test of predictive power, all five BELT hyperensembles recover set-aside measurements not used in the fitting and report accu- rate errors, even when starting from highly inaccurate simulations. BELTs principled fxramework thus enables practical predictions for complex biomolecular systems from discordant simulations and sparse data.

Biophysics

DISEASES: Text mining and data integration of disease–gene associations

Text mining is a flexible technology that can be applied to numerous different tasks in biology and medicine. We present a system for extracting disease-gene associations from biomedical abstracts. The system consists of a highly efficient dictionary-based tagger for named entity recognition of human genes and diseases, which we combine with a scoring scheme that takes into account co-occurrences both within and between sentences. We show that this approach is able to extract half of all manually curated associations with a false positive rate of only 0.16%. Nonetheless, text mining should not stand alone, but be combined with other types of evidence. For this reason, we have developed the DISEASES resource, which integrates the results from text mining with manually curated disease-gene associations, cancer mutation data, and genome-wide association studies from existing databases. The DISEASES resource is accessible through a user-friendly web interface at http://diseases.jensenlab.org/, where the text-mining software and all associations are also freely available for download.

Bioinformatics