Search bioRxivSearch

Biology subjects

Cantini, L.

Publications and source records attributed to Cantini, L..

5 recordsLinked to original sources

Stabilized Independent Component Analysis outperforms other methods in finding reproducible signals in tumoral transcriptomes

MotivationMatrix factorization methods are widely exploited in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). Applying such methods to similar independent datasets should yield reproducible inter-series outputs, though it was never demonstrated yet.\n\nResultsWe systematically test state-of-art methods of matrix factorization on several transcriptomic datasets of the same cancer type. Inspired by concepts of evolutionary bioinformatics, we design a new framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the methods reproducibility. We show that a particular protocol of application of Independent Component Analysis (ICA), accompanied by a stabilisation procedure, leads to a significant increase in the inter-series output reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other state-of-art methods. We developed a user-friendly tool BIODICA for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent publicly available transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors that can be associated to biological processes or to technological artefacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping.\n\nAvailabilityThe BIODICA tool is available from https://github.com/LabBandSB/BIODICA.\n\nContactlaura.cantini@curie.fr and andrei.zinovyev@curie.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Identification of microRNA clusters cooperatively acting on Epithelial to Mesenchymal Transition in Triple Negative Breast Cancer

MicroRNAs play important roles in many biological processes. Their aberrant expression can have oncogenic or tumor suppressor function directly participating to carcinogenesis, malignant transformation, invasiveness and metastasis. Indeed, miRNA profiles can distinguish not only between normal and cancerous tissue but they can also successfully classify different subtypes of a particular cancer. Here, we focus on a particular class of transcripts encoding polycistronic miRNA genes that yields multiple miRNA components. We describe clustered MiRNA Master Regulator Analysis (ClustMMRA), a fully redesigned release of the MMRA computational pipeline (MiRNA Master Regulator Analysis), developed to search for clustered miRNAs potentially driving cancer molecular subtyping. Genomically clustered miRNAs are frequently co-expressed to target different components of pro-tumorigenic signalling pathways. By applying ClustMMRA to breast cancer patient data, we identified key miRNA clusters driving the phenotype of different tumor subgroups. The pipeline was applied to two independent breast cancer datasets, providing statistically concordant results between the two analysis. We validated in cell lines the miR-199/miR-214 as a novel cluster of miRNAs promoting the triple negative subtype phenotype through its control of proliferation and EMT.

systems biology

Moonlight: a tool for biological interpretation and driver genes discovery

Cancer is a complex and heterogeneous disease. It is crucial to identify the key driver genes and their role in cancer mechanisms with attention to different cancer stages, types or subtypes. Cancer driver genes are elusive and their discovery is complicated by the fact that the same gene can play a diverse role in different contexts. Key biological processes, such as cell proliferation and cell death, have been linked to cancer progression. Thus, in principle, they can be exploited to classify the cancer genes and unveil their role. Here, we present a new method, Moonlight, that exploit expression data to classify cancer genes. Moonlight relies on the integration of functional enrichment analysis, gene regulatory networks and upstream regulator analysis from expression data to score the importance of biological cancer-related processes taking into account either the inter- or intra-tumor heterogeneity. We then employed these scores to predict if each gene acts as a tumor suppressor gene (TSG) or as an oncogene (OCG). Our methodology also allow to predict genes with dual role, i.e. the moonlight genes (TSG in one cancer type or stage and OCG in another), as well as to elucidate the underlying biological processes. Availability: https://bioconductor.org/packages/MoonlightR & https://github.com/ibsquare/MoonlightR/

bioinformatics

Determining the optimal number of independent components for reproducible transcriptomic data analysis

BackgroundIndependent Component Analysis (ICA) is a method that models gene expression data as an action of a set of statistically independent hidden factors. The output of ICA depends on a fundamental parameter: the number of components (factors) to compute. The optimal choice of this parameter, related to determining the effective data dimension, remains an open question in the application of blind source separation techniques to transcriptomic data.\n\nResultsHere we address the question of optimizing the number of statistically independent components in the analysis of transcriptomic data for reproducibility of the components in multiple runs of ICA (within the same or within varying effective dimensions) and in multiple independent datasets. To this end, we introduce ranking of independent components based on their stability in multiple ICA computation runs and define a distinguished number of components (Most Stable Transcriptome Dimension, MSTD) corresponding to the point of the qualitative change of the stability profile. Based on a large body of data, we demonstrate that a sufficient number of dimensions is required for biological interpretability of the ICA decomposition and that the most stable components with ranks below MSTD have more chances to be reproduced in independent studies compared to the less stable ones. At the same time, we show that a transcriptomics dataset can be reduced to a relatively high number of dimensions without losing the interpretability of ICA, even though higher dimensions give rise to components driven by small gene sets.\n\nConclusionsWe suggest a protocol of ICA application to transcriptomics data with a possibility of prioritizing components with respect to their reproducibility that strengthens the biological interpretation. Computing too few components (much less than MSTD) is not optimal for interpretability of the results. The components ranked within MSTD range have more chances to be reproduced in independent studies.

bioinformatics

Classification Of Gene Signatures For Their Information Value And Functional Redundancy

Large collections of gene signatures play a pivotal role in interpreting results of omics data analysis but suffer from compositional (large overlap) and functional (redundant read-outs) redundancy, and many gene signatures rarely pop-up in statistical tests. Based on pan-cancer data analysis, here we define a restricted set of 962 so called informative signatures and demonstrate that they have more chances to appear highly enriched in cancer biology studies. We show that the majority of informative signatures conserve their weights for the composing genes (eigengenes) from one cancer type to another. We construct InfoSigMap, an interactive online map showing the structure of compositional and functional redundancies between informative signatures and charting the territories of biological functions accessible through transcriptomic studies. InfoSigMap can be used to visualize in one insightful picture the results of comparative omics data analyses and suggests reconsidering existing annotations of certain reference gene set groups.

systems biology