Search bioRxivSearch

Biology subjects

Casey S Greene

Publications and source records attributed to Casey S Greene.

10 recordsLinked to original sources

System-wide automatic extraction of functional signatures in Pseudomonas aeruginosa with eADAGE

Cross experiment comparisons in public data compendia are challenged by unmatched conditions and technical noise. The ADAGE method, which performs unsupervised integration with neural networks, can effectively identify biological patterns, but because ADAGE models, like many neural networks, are over-parameterized, different ADAGE models perform equally well. To enhance model robustness and better build signatures consistent with biological pathways, we developed an ensemble ADAGE (eADAGE) that integrated stable signatures across models. We applied eADAGE to a Pseudomonas aeruginosa compendium containing experiments performed in 78 media. eADAGE revealed a phosphate starvation response controlled by PhoB. While we expected PhoB activity in limiting phosphate conditions, our analyses found PhoB activity in other media with moderate phosphate and predicted that a second stimulus provided by the sensor kinase, KinB, is required for PhoB activation in this setting. We validated this relationship using both targeted and unbiased genetic approaches. eADAGE, which captures stable biological patterns, enables cross-experiment comparisons that can highlight measured but undiscovered relationships.

Systems Biology

A machine learning classifier trained on cancer transcriptomes detects NF1 inactivation signal in glioblastoma

Background: We have identified molecules that exhibit synthetic lethality in cells with loss of the neurofibromin 1 (NF1) tumor suppressor gene. However, recognizing tumors that have inactivation of the NF1 tumor suppressor function is challenging because the loss may occur via mechanisms that do not involve mutation of the genomic locus. Degradation of the NF1 protein, independent of NF1 mutation status, photocopies inactivating mutations to drive tumors in human glioma cell lines. NF1 inactivation may alter the transcriptional landscape of a tumor and allow a machine learning classifier to detect which tumors will benefit from synthetic lethal molecules.\n\nResults: We developed a strategy to predict tumors with low NF1 activity and hence tumors that may respond to treatments that target cells lacking NF1. Using RNAseq data from The Cancer Genome Atlas (TCGA), we trained an ensemble of 500 logistic regression classifiers that integrates mutation status with whole transcriptomes to predict NF1 inactivation in glioblastoma (GBM). On TCGA data, the classifier detected NF1 mutated tumors (test set area under the receiver operating characteristic curve (AUROC) mean = 0.77, 95% quantile = 0.53 - 0.95) over 50 random initializations. On RNA-Seq data transformed into the space of gene expression microarrays, this method produced a classifier with similar performance (test set AUROC mean = 0.77, 95% quantile = 0.53 - 0.96). We applied our ensemble classifier trained on the transformed TCGA data to a microarray validation set of 12 samples with matched RNA and NF1 protein-level measurements. The classifiers NF1 score was associated with NF1 protein concentration in these samples.\n\nConclusions: We demonstrate that TCGA can be used to train accurate predictors of NF1 inactivation in GBM. The ensemble classifier performed well for samples with very high or very low NF1 protein concentrations but had mixed performance in samples with intermediate NF1 concentrations. Nevertheless, high-performing and validated predictors have the potential to be paired with targeted therapies and personalized medicine.

Genomics

Integrative networks illuminate biological factors underlying gene-disease associations

A.Integrative networks combine multiple layers of biological data into a model of how genes work together to carry out cellular processes. Such networks become more valuable as they become more context specific, for example, by capturing how genes work together in a certain tissue or cell type. Once constructed, these networks provide the means to identify broad biological patterns underlying genes associated with complex traits and diseases. In this review, we discuss the different types of integrative networks that currently exist and how such networks that encompass multiple biological layers are constructed. We highlight how specificity can be incorporated into the reconstruction of different types of biomolecular interactions between genes, using tissue-specificity as a motivating example. We identify examples of cases where networks have been applied to study human diseases and discuss opportunities for new applications.

Bioinformatics

Reproducible Computational Workflows with Continuous Analysis

Reproducing experiments is vital to science. Being able to replicate, validate and extend previous work also speeds new research projects. Reproducing computational biology experiments, which are scripted, should be straightforward. But reproducing such work remains challenging and time consuming. In the ideal world we would be able to quickly and easily rewind to the precise computing environment where results were generated. We would then be able to reproduce the original analysis or perform new analyses. We introduce a process termed \"continuous analysis\" which provides inherent reproducibility to computational research at a minimal cost to the researcher. Continuous analysis combines Docker, a container service similar to virtual machines, with continuous integration, a popular software development technique, to automatically re-run computational analysis whenever relevant changes are made to the source code. This allows results to be reproduced quickly, accurately and without needing to contact the original authors. Continuous analysis also provides an audit trail for analyses that use data with sharing restrictions. This allows reviewers, editors, and readers to verify reproducibility without manually downloading and rerunning any code. Example configurations are available at our online repository (https://github.com/greenelab/continuous_analysis).

Bioinformatics

Tribe: The collaborative platform for reproducible web-based analysis of gene sets

BackgroundThe adoption of new bioinformatics webservers provides biological researchers with new analytical opportunities but also raises workflow challenges. These challenges include sharing collections of genes with collaborators, translating gene identifiers to the most appropriate nomenclature for each server, tracking these collections across multiple analysis tools and webservers, and maintaining effective records of the genes used in each analysis.\n\nDescriptionIn this paper, we present the Tribe webserver (available at https://tribe.greenelab.com), which addresses these challenges in order to make multi-server workflows seamless and reproducible. This allows users to create analysis pipelines that use their own sets of genes in combinations of specialized data mining webservers and tools while seamlessly maintaining gene set version control. Tribes web interface facilitates collaborative editing: users can share with collaborators, who can then view, download, and edit these collections. Tribes fully-featured API allows users to interact with Tribe programmatically if desired. Tribe implements the OAuth 2.0 standard as well as gene identifier mapping, which facilitates its integration into existing servers. Access to Tribes resources is facilitated by an easy-to-install Python application called tribe-client. We provide Tribe and tribe-client under a permissive open-source license to encourage others to download the source code and set up a local instance or to extend its capabilities.\n\nConclusionsThe Tribe webserver addresses challenges that have made reproducible multi-webserver workflows difficult to implement until now. It is open source, has a user-friendly web interface, and provides a means for researchers to perform reproducible gene set based analyses seamlessly across webservers and command line tools.

Bioinformatics

Pathway and network-based strategies to translate genetic discoveries into effective therapies

One way to design a drug is to attempt to phenocopy a genetic variant that is known to have the desired effect. In general, drugs that are supported by genetic associations progress further in the development pipeline. However, the number of associations that are candidates for development into drugs is limited because many associations are in noncoding regions or difficult to target genes. Approaches that overlay information from pathway databases or biological networks can expand the potential target list. In cases where the initial variant is not targetable or there is no variant with the desired effect, this may reveal new means to target a disease. In this review we discuss recent examples in the domain of pathway and network-based drug repositioning from genetic associations. We highlight important caveats and challenges for the field, and we discuss opportunities for further development.

Genetics

Semi-Supervised Learning of the Electronic Health Record for Phenotype Stratification

Patient interactions with health care providers result in entries to electronic health records (EHRs). EHRs were built for clinical and billing purposes but contain many data points about an individual. Mining these records provides opportunities to extract electronic phenotypes, which can be paired with genetic data to identify genes underlying common human diseases. This task remains challenging: high quality phenotyping is costly and requires physician review; many fields in the records are sparsely filled; and our definitions of diseases are continuing to improve over time. Here we develop and evaluate a semi-supervised learning method for EHR phenotype extraction using denoising autoencoders for phenotype stratification. By combining denoising autoencoders with random forests we find classification improvements across multiple simulation models and improved survival prediction in ALS clinical trial data. This is particularly evident in cases where only a small number of patients have high quality phenotypes, a common scenario in EHR-based research. Denoising autoencoders perform dimensionality reduction enabling visualization and clustering for the discovery of new subtypes of disease. This method represents a promising approach to clarify disease subtypes and improve genotype-phenotype association studies that leverage EHRs.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=81 SRC=\"FIGDIR/small/039800_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (31K):\norg.highwire.dtl.DTLVardef@567074org.highwire.dtl.DTLVardef@f0e45corg.highwire.dtl.DTLVardef@1207659org.highwire.dtl.DTLVardef@39f8b8_HPS_FORMAT_FIGEXP M_FIG C_FIG HIGHLIGHTSO_LIDenoising autoencoders (DAs) can model electronic health records.\nC_LIO_LISemi-supervised learning with DAs improves ALS patient survival predictions.\nC_LIO_LIDAs improve patient cluster visualization through dimensionality reduction.\nC_LI

Bioinformatics

A Novel Multi-network Approach Reveals Tissue-specific Cellular Modulators of Fibrosis in Systemic Sclerosis, Pulmonary Fibrosis and Pulmonary Arterial Hypertension

We have used integrative genomics to determine if a common molecular mechanism underlies different clinical manifestations in systemic sclerosis (SSc), and the related conditions pulmonary fibrosis (PF) and pulmonary arterial hypertension (PAH). We identified a common pathogenic gene expression signature - an immune-fibrotic axis-indicative of pro-fibrotic macrophages in multiple affected tissues (skin, lung, esophagus and PBMCs) of SSc, PF, and PAH. We used this disease-associated signature to query tissue-specific functional genomic networks. This allowed us to identify common and tissue-specific pathology of SSc and related conditions. We rigorously contrasted the lung- and skin-specific gene-gene interaction networks to identify a distinct lung resident macrophage signature associated with lipid stimulation and alternative activation. In keeping with our network results, we find distinct macrophages alternative activation transcriptional programs in SSc-PF lung and in the skin of patients with an \"inflammatory\" SSc gene expression signature. Our results suggest that the innate immune system is central to SSc disease processes, but that subtle distinctions exist between tissues. Our approach provides a framework for examining molecular signatures of disease in fibrosis and autoimmune diseases and for leveraging publicly available data to understand common and tissue-specific disease processes in complex human diseases.

Systems Biology

ADAGE analysis of publicly available gene expression data collections illuminates Pseudomonas aeruginosa-host interactions

The growth in genome-scale assays of gene expression for different species in publicly available databases presents new opportunities for computational methods that aid in hypothesis generation and biological interpretation of these data. Here, we present an unsupervised machine-learning approach, ADAGE (Analysis using Denoising Autoencoders of Gene Expression) and apply it to the interpretation of all of the publicly available gene expression data for Pseudomonas aeruginosa, an important opportunistic bacterial pathogen. In post-hoc positive control analyses using curated knowledge, the P. aeruginosa ADAGE model found that co-operonic genes often participated in similar processes and accurately predicted which genes had similar functions. By analyzing newly generated data and previously published microarray and RNA-seq data, the ADAGE model identified gene expression differences between strains, modeled the cellular response to low oxygen, and predicted the involvement of biological processes despite low level expression differences in directly involved genes. Comparison of ADAGE with PCA and ICA revealed that ADAGE extracts distinct signals. We provide the ADAGE model with analysis of all publicly available P. aeruginosa GeneChip experiments, and we provide open source code for use in other species and settings.

Bioinformatics

Comprehensive cross-population analysis of high-grade serous ovarian cancer supports no more than three subtypes

BackgroundThree to four gene expression-based subtypes of high-grade serous ovarian cancer (HGSC) have been previously reported. We sought to systematically determine the similarity of HGSC subtypes between populations.\n\nMethodsWe independently clustered (k = 3 and k = 4) five publicly-available HGSC mRNA expression datasets with >130 tumors using k-means and non-negative matrix factorization. Within each population, we summarized differential expression patterns for each cluster as moderated t statistic vectors using Significance Analysis of Microarrays. We calculated Pearsons correlations of these vectors to determine similarities and differences in expression patterns between clusters. We defined syn-clusters (SC) as sets of clusters that were strongly correlated across populations, and associated their expression patterns with biological pathways using geneset overrepresentation analyses.\n\nResultsAcross populations, for k = 3, moderated t score correlations for clusters 1, 2 and 3, respectively, ranged between 0.77-0.85, 0.80-0.90, and 0.65-0.77. For k = 4, correlations for clusters 1-4, respectively, ranged between 0.77-0.85, 0.83-0.89, 0.51-0.76, and 0.61-0.75. Within populations, comparing analogous clusters (k = 3 versus k = 4), correlations were high for clusters 1 and 2 (0.91-1.00), but were lower for cluster 3 (0.22-0.80). Results are similar using non-negative matrix factorization. SC1 corresponds to previously-reported mesenchymal-like, SC2 to proliferative-like, SC3 to immunoreactive-like, and SC4 to differentiated-like subtypes.\n\nConclusionsThe mesenchymal-like and proliferative-like subtypes are remarkably consistent across populations and could be uniquely targeted for treatment. The other two previously described subtypes are considerably less robust, and since cross-population comparison reveals that k = 3 and k = 4 are both consistent with our results, they may not represent clear subtypes.

Cancer Biology