Search bioRxivSearch

bioRxiv · 10.1101/043257

Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer

Abstract

BackgroundDe novo inference of clinically relevant gene function relationships from tumor RNA-seq remains a challenging task. Current methods typically either partition patient samples into a few subtypes or rely upon analysis of pairwise gene correlations (co-expression) that will miss some groups in noisy data. Leveraging higher dimensional information can be expected to increase the power to discern targetable pathways, but this is commonly thought to be an intractable computational problem.\n\nMethodsIn this work we adapt a recently developed machine learning algorithm, CorEx, that efficiently optimizes over multivariate mutual information for sensitive detection of complex gene relationships. The algorithm can be iteratively applied to generate a hierarchy of latent factors. Patients are stratified relative to each factor and combinatoric survival analyses are performed and interpreted in the context of biological function annotations and protein network interactions that might be utilized to match patients to multiple therapies.\n\nResultsAnalysis of ovarian tumor RNA-seq samples demonstrates the algorithms power to infer well over one hundred biologically interpretable gene cohorts, several times more than standard methods such as hierarchical clustering and k-means. The CorEx factor hierarchy is also informative, with related but distinct gene clusters grouped by upper nodes. Some latent factors correlate with patient survival, including one for a pathway connected with the epithelial-mesenchymal transition in breast cancer that is regulated by a potentially druggable microRNA. Further, combinations of factors lead to a synergistic survival advantage in some cases.\n\nConclusionsIn contrast to studies that attempt to partition patients into a small number of subtypes (typically 4 or fewer) for treatment purposes, our approach utilizes subgroup information for combinatoric transcriptional phenotyping. Considering only the 66 gene expression groups that are both found to have significant Gene Ontology enrichment and are small enough to indicate specific drug targets implies a computational phenotype for ovarian cancer that allows for 366 possible patient profiles, enabling truly personalized treatment. The findings here demonstrate a new technique that sheds light on the complexity of gene expression dependencies in tumors and could eventually enable the use of patient RNA-seq profiles for selection of personalized and effective cancer treatments.

Source connections

Explore related subjects

Keep this discovery

BibTeXRIS

Shirley Pepke, Greg Ver Steeg. 2016-03-11. Comprehensive discovery of subsample gene expression components by information explanation: therapeutic implications in cancer. https://doi.org/10.1101/043257

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Single Cell Phenotyping Reveals Heterogeneity among Haematopoietic Stem Cells Following Infection

The haematopoietic stem cell (HSC) niche provides essential micro-environmental cues for the production and maintenance of HSCs within the bone marrow. During inflammation, haematopoietic dynamics are perturbed, but it is not known whether changes to the HSC-niche interaction occur as a result. We visualise HSCs directly in vivo, enabling detailed analysis of the 3D niche dynamics and migration patterns in murine bone marrow following Trichinella spiralis infection. Spatial statistical analysis of these HSC trajectories reveals two distinct modes of HSC behaviour: (i) a pattern of revisiting previously explored space, and (ii) a pattern of exploring new space. Whereas HSCs from control donors predominantly follow pattern (i), those from infected mice adopt both strategies. Using detailed computational analyses of cell migration tracks and life-history theory, we show that the increased motility of HSCs following infection can, perhaps counterintuitively, enable mice to cope better in deteriorating HSC-niche micro-environments following infection.\n\nAuthor SummaryHaematopoietic stem cells reside in the bone marrow where they are crucially maintained by an incompletely-determined set of niche factors. Recently it has been shown that chronic infection profoundly affects haematopoiesis by exhausting stem cell function, but these changes have not yet been resolved at the single cell level. Here we show that the stem cell-niche interactions triggered by infection are heterogeneous whereby cells exhibit different behavioural patterns: for some, movement is highly restricted, while others explore much larger regions of space over time. Overall, cells from infected mice display higher levels of persistence. This can be thought of as a search strategy: during infection the signals passed between stem cells and the niche may be blocked or inhibited. Resultantly, stem cells must choose to either cling on, or to leave in search of a better environment. The heterogeneity that these cells display has immediate consequences for translational therapies involving bone marrow transplant, and the effects that infection might have on these procedures.

Systems Biology

Analysis of noise mechanisms in cell size control

At the single-cell level, noise features in multiple ways through the inherent stochasticity of biomolecular processes, random partitioning of resources at division, and fluctuations in cellular growth rates. How these diverse noise mechanisms combine to drive variations in cell size within an isoclonal population is not well understood. To address this problem, we systematically investigate the contributions of different noise sources in well-known paradigms of cell-size control, such as the adder (division occurs after adding a fixed size from birth) and the sizer (division occurs upon reaching a size threshold). Analysis reveals that variance in cell size is most sensitive to errors in partitioning of volume among daughter cells, and not surprisingly, this process is well regulated among microbes. Moreover, depending on the dominant noise mechanism, different size control strategies (or a combination of them) provide efficient buffering of intercellular size variations. We further explore mixer models of size control, where a timer phase precedes/follows an adder, as has been proposed in Caulobacter crescentus. While mixing a timer with an adder can sometimes attenuate size variations, it invariably leads to higher-order moments growing unboundedly over time. This results in the cell size following a power-law distribution with an exponent that is inversely dependent on the noise in the timer phase. Consistent with theory, we find evidence of power-law statistics in the tail of C. crescentus cell-size distribution, but there is a huge discrepancy in the power-law exponent as estimated from data and theory. However, the discrepancy is removed after data reveals that the size added by individual newborns from birth to division itself exhibits power-law statistics. Taken together, this study provides key insights into the role of noise mechanisms in size homeostasis, and suggests an inextricable link between timer-based models of size control and heavy-tailed cell size distributions.

Systems Biology

A mathematical approach for secondary structure analysis can provide an eyehole to the RNA world

The RNA pseudoknot is a conserved secondary structure encountered in a number of ribozymes, which assume a central role in the RNA world hypothesis. However, RNA folding algorithms could not predict pseudoknots until recently. Analytic combinatorics - a newly arisen mathematical field - has introduced a way of enumerating different RNA configurations and quantifying RNA pseudoknot structure robustness and evolvability, two features that drive their molecular evolution. I will present a mathematicians viewpoint of RNA secondary structures, and explain how analytic combinatorics applied on RNA sequence to structure maps can represent a valuable tool for understanding RNA secondary structure evolution. Analytic combinatorics can be implemented for the optimization of RNA secondary structure prediction algorithms, the derivation of molecular evolution mathematical models, as well as in a number of biotechnological applications, such as biosensors, riboswitches etc. Moreover, it showcases how the integration of biology and mathematics can provide a different viewpoint into the RNA world.

Systems Biology