Search bioRxivSearch

bioRxiv · 10.1101/048215

Extracting a low-dimensional description of multiple gene expression datasets reveals a potential driver for tumor-associated stroma in ovarian cancer

Abstract

BackgroundDiscovering patient subtypes and molecular drivers of a subtype are difficult and driving problems underlying most modern disease expression studies collected across patient populations. Expression patterns conserved across multiple expression datasets from independent disease studies are likely to represent important molecular events underlying the disease.\n\nMethodsWe present the INSPIRE (INferring Shared modules from multiPle gene expREssion datasets) method to infer highly coherent and robust modules of co-expressed genes and the dependencies among the modules from multiple expression datasets. Focusing on inferring modules and their dependencies conserved across multiple expression datasets is important for several reasons. First, using multiple datasets will increase the power to detect robust and relevant patterns (modules and dependencies among modules). Second, INSPIRE enables the use of multiple datasets that contain different sets of genes due to, e.g., the difference in microarray platforms. Many methods designed for expression data analysis cannot integrate multiple datasets with variable discrepancy to infer a single combined model, whereas INSPIRE can naturally model the dependencies among the modules even when a large proportion of genes are not observed on a certain platform.\n\nResultsWe evaluated INSPIRE on synthetically generated datasets with known underlying network structure among modules, and gene expression datasets from multiple ovarian cancer studies. We show that the model learned by INSPIRE can explain unseen data better and can reveal prior knowledge on gene functions more accurately than alternative methods. We demonstrate that applying INSPIRE to nine ovarian cancer datasets leads to the identification of a new marker and potential molecular driver of tumor-associated stroma - HOPX. We also demonstrate that the HOPXmodule strongly overlaps with the genes defining the mesenchymal patient subtype identified in The Cancer Genome Atlas (TCGA) ovarian cancer data. We provide evidence for a previously unknown molecular basis of tumor resectability efficacy involving tumor-associated mesenchymal stem cells represented by HOPX.\n\nConclusionsINSPIRE extracts a low-dimensional description from multiple gene expression data, which consists of modules and their dependencies. The discovery of a new tumor-associated stroma marker, HOPX, and its module suggests a previously unknown mechanism underlying tumor-associated stroma.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Safiye Celik, Benjamin A Logsdon, Stephanie Battle, Charles W Drescher, Mara Rendi, David Hawkins, Su-In Lee. 2016-04-13. Extracting a low-dimensional description of multiple gene expression datasets reveals a potential driver for tumor-associated stroma in ovarian cancer. https://doi.org/10.1101/048215

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Single Cell Phenotyping Reveals Heterogeneity among Haematopoietic Stem Cells Following Infection

The haematopoietic stem cell (HSC) niche provides essential micro-environmental cues for the production and maintenance of HSCs within the bone marrow. During inflammation, haematopoietic dynamics are perturbed, but it is not known whether changes to the HSC-niche interaction occur as a result. We visualise HSCs directly in vivo, enabling detailed analysis of the 3D niche dynamics and migration patterns in murine bone marrow following Trichinella spiralis infection. Spatial statistical analysis of these HSC trajectories reveals two distinct modes of HSC behaviour: (i) a pattern of revisiting previously explored space, and (ii) a pattern of exploring new space. Whereas HSCs from control donors predominantly follow pattern (i), those from infected mice adopt both strategies. Using detailed computational analyses of cell migration tracks and life-history theory, we show that the increased motility of HSCs following infection can, perhaps counterintuitively, enable mice to cope better in deteriorating HSC-niche micro-environments following infection.\n\nAuthor SummaryHaematopoietic stem cells reside in the bone marrow where they are crucially maintained by an incompletely-determined set of niche factors. Recently it has been shown that chronic infection profoundly affects haematopoiesis by exhausting stem cell function, but these changes have not yet been resolved at the single cell level. Here we show that the stem cell-niche interactions triggered by infection are heterogeneous whereby cells exhibit different behavioural patterns: for some, movement is highly restricted, while others explore much larger regions of space over time. Overall, cells from infected mice display higher levels of persistence. This can be thought of as a search strategy: during infection the signals passed between stem cells and the niche may be blocked or inhibited. Resultantly, stem cells must choose to either cling on, or to leave in search of a better environment. The heterogeneity that these cells display has immediate consequences for translational therapies involving bone marrow transplant, and the effects that infection might have on these procedures.

Systems Biology

Analysis of noise mechanisms in cell size control

At the single-cell level, noise features in multiple ways through the inherent stochasticity of biomolecular processes, random partitioning of resources at division, and fluctuations in cellular growth rates. How these diverse noise mechanisms combine to drive variations in cell size within an isoclonal population is not well understood. To address this problem, we systematically investigate the contributions of different noise sources in well-known paradigms of cell-size control, such as the adder (division occurs after adding a fixed size from birth) and the sizer (division occurs upon reaching a size threshold). Analysis reveals that variance in cell size is most sensitive to errors in partitioning of volume among daughter cells, and not surprisingly, this process is well regulated among microbes. Moreover, depending on the dominant noise mechanism, different size control strategies (or a combination of them) provide efficient buffering of intercellular size variations. We further explore mixer models of size control, where a timer phase precedes/follows an adder, as has been proposed in Caulobacter crescentus. While mixing a timer with an adder can sometimes attenuate size variations, it invariably leads to higher-order moments growing unboundedly over time. This results in the cell size following a power-law distribution with an exponent that is inversely dependent on the noise in the timer phase. Consistent with theory, we find evidence of power-law statistics in the tail of C. crescentus cell-size distribution, but there is a huge discrepancy in the power-law exponent as estimated from data and theory. However, the discrepancy is removed after data reveals that the size added by individual newborns from birth to division itself exhibits power-law statistics. Taken together, this study provides key insights into the role of noise mechanisms in size homeostasis, and suggests an inextricable link between timer-based models of size control and heavy-tailed cell size distributions.

Systems Biology

A mathematical approach for secondary structure analysis can provide an eyehole to the RNA world

The RNA pseudoknot is a conserved secondary structure encountered in a number of ribozymes, which assume a central role in the RNA world hypothesis. However, RNA folding algorithms could not predict pseudoknots until recently. Analytic combinatorics - a newly arisen mathematical field - has introduced a way of enumerating different RNA configurations and quantifying RNA pseudoknot structure robustness and evolvability, two features that drive their molecular evolution. I will present a mathematicians viewpoint of RNA secondary structures, and explain how analytic combinatorics applied on RNA sequence to structure maps can represent a valuable tool for understanding RNA secondary structure evolution. Analytic combinatorics can be implemented for the optimization of RNA secondary structure prediction algorithms, the derivation of molecular evolution mathematical models, as well as in a number of biotechnological applications, such as biosensors, riboswitches etc. Moreover, it showcases how the integration of biology and mathematics can provide a different viewpoint into the RNA world.

Systems Biology