Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Cancer Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Personalized Regression Enables Sample-Specific Pan-Cancer Analysis

In many applications, inter-sample heterogeneity is crucial to understanding the complex biological processes under study. For example, in genomic analysis of cancers, each patient in a cohort may have a different driver mutation, making it difficult or impossible to identify causal mutations from an averaged view of the entire cohort. Unfortunately, many traditional methods for genomic analysis seek to estimate a single model which is shared by all samples in a population, ignoring this inter-sample heterogeneity entirely. In order to better understand patient heterogeneity, it is necessary to develop practical, personalized statistical models. To uncover this inter-sample heterogeneity, we propose a novel regularizer for achieving patient-specific personalized estimation. This regularizer operates by learning two latent distance metrics - one between personalized parameters and one between clinical covariates - and attempting to match the induced distances as closely as possible. Crucially, we do not assume these distance metrics are already known. Instead, we allow the data to dictate the structure of these latent distance metrics. Finally, we apply our method to learn patient-specific, interpretable models for a pan-cancer gene expression dataset containing samples from more than 30 distinct cancer types and find strong evidence of personalization effects between cancer types as well as between individuals. Our analysis uncovers sample-specific aberrations that are overlooked by population level methods, suggesting a promising new path for precision analysis of complex diseases such as cancer.

bioinformatics

Comparing cancer cell lines and tumor samples by genomic profiles

Cancer cell lines are often used in laboratory experiments as models of tumors, although they can have substantially different genetic and epigenetic profiles compared to tumors. We have developed a general computational method - TumorComparer - to systematically quantify similarities and differences between tumor material when detailed genetic and molecular profiles are available. The comparisons can be flexibly tailored to a particular biological question by placing a higher weight on functional alterations of interest ( weighted similarity). In a first pan-cancer application, we have compared 260 cell lines from the Cancer Cell Line Encyclopaedia (CCLE) and 1914 tumors of six different cancer types from The Cancer Genome Atlas (TCGA), using weights to emphasize genomic alterations that frequently recur in tumors. We report the potential suitability of particular cell lines as tumor models and identify apparently unsuitable outlier cell lines, some of which are in wide use, for each of the six cancer types. In future, this weighted similarity method may be generalized for use in a clinical setting to compare patient profiles consisting of genomic patterns combined with clinical attributes, such as diagnosis, treatment and response to therapy.

Cancer Biology

Discovery and optimization of piperazine-1-thiourea-based human phosphoglycerate dehydrogenase inhibitors

Proliferating cells, including cancer cells, obtain serine both exogenously and via the metabolism of glucose. By catalyzing the first, rate-limiting step in the synthesis of serine from glucose, phosphoglycerate dehydrogenase (PHGDH) controls flux through the biosynthetic pathway for this important amino acid and represents a putative target in oncology. To discover inhibitors of PHGDH, a coupled biochemical assay was developed and optimized to enable high-throughput screening for inhibitors of human PHGDH. Feedback inhibition was minimized by coupling PHGDH activity to two downstream enzymes (PSAT1 and PSPH), providing a significant improvement in enzymatic turnover. Further coupling of NADH to a diaphorase/resazurin system enabled a red-shifted detection readout, minimizing interference due to compound autofluorescence. With this protocol, over 400,000 small molecules were screened for PHGDH inhibition, and following hit validation and triage work, a piperazine-1-thiourea was identified. Following rounds of medicinal chemistry and SAR exploration, two probes (NCT-502 and NCT-503) were identified. These molecules demonstrated improved target activity and encouraging ADME properties, enabling both in vitro and in vivo assessment of the biological importance of PHGDH, and its role in the fate of serine in PHGDH-dependent cancer cells.

cancer biology

IMO-HIP 2015 Report: An Evolutionary Game Theory Approach to evolutionary-enlightened application of chemotherapy in bone metastatic prostate cancer

Prostate cancer metastasis to the bone is predominantly lethal and results from the ability of successful metastatic prostate cancer cells to co-opt microenvironmental cells and processes involved in bone remodelling. Understanding how the interactions between tumour and stromal cells determine successful metastases and how metastatic tumours respond to treatment is an emergent process that is hard to asses biologically and thus can benefit from mathematical models. In this work we describe a mathematical model of bone remodelling and the establishment of a prostate cancer metastasis in the bone using evolutionary game theory. We have mathematically recapitulated the current paradigm of a vicious cycle driving the tumor growth and we have used this tool to investigate the key interactions between the tumour and the bone stroma. Crucially, the model sheds light on the role that the interactions of heterogeneous tumor cells with the bone microenvironment have in the treatment of cancer. Our results show that resistant populations naturally become dominant in the metastases under a number treatment schemes and that schedules designed by an evolutionary game theory approach could be used to better control the tumour and the associated bone growth than the current standard of care.

Cancer Biology

Molecular characterization of breast and lung tumors by integration of multiple data types with sparse-factor analysis

Effective cancer treatment is crucially dependent on the identification of the biological processes that drive a tumor. However, multiple processes may be active simultaneously in a tumor. Clustering is inherently unsuitable to this task as it assigns a tumor to a single cluster. In addition, the wide availability of multiple data types per tumor provides the opportunity to profile the processes driving a tumor more comprehensively.\n\nHere we introduce Functional Sparse-Factor Analysis (funcSFA) to address these challenges. FuncSFA integrates multiple data types to define a lower dimensional space capturing the relevant variation. A tailor-made module associates biological processes with these factors. FuncSFA is inspired by iCluster, which we improve in several key aspects. First, we increase the convergence efficiency significantly, allowing the analysis of multiple molecular datasets that have not been pre-matched to contain only concordant features. Second, FuncSFA does not assign tumors to discrete clusters, but identifies the dominant driver processes active in each tumor. This is achieved by a regression of the factors on the RNA expression data followed by a functional enrichment analysis and manual curation step.\n\nWe apply FuncSFA to the TCGA breast and lung datasets. We identify EMT and Immune processes common to both cancer types. In the breast cancer dataset we recover the known intrinsic subtypes and identify additional processes. These include immune infiltration and EMT, and processes driven by copy number gains on the 8q chromosome arm. In lung cancer we recover the major types (adenocarcinoma and squamous cell carcinoma) and processes active in both of these types. These include EMT, two immune processes, and the activity of the NFE2L2 transcription factor.\n\nIn summary, FuncSFA is a robust method to perform discovery of key driver processes in a collection of tumors through unsupervised integration of multiple molecular data types and functional annotation.\n\nAuthor SummaryIn order to select effective cancer treatment, we need to determine which biological processes are active in a tumor. To this end, tumors have been quantified by high dimensional molecular measurements such as RNA sequencing and DNA copy number profiling. In order to support decision making, these measurements need to be condensed into interpretable summaries. Such summaries can be made interpretable by connecting them to biological processes.\n\nBiological process activity is continuous and multiple biological processes are taking place in a single tumor. Therefore, the biological processes associated with a tumor are misrepresented by clustering, which tries to put every tumor in a single cluster. In the method introduced in this paper (funcSFA), molecular measurements are summarized into a small number factors. A factor is a continuous value per tumor that aims to represent the activity of a biological process.\n\nWhen applied to breast and lung cancer, funcSFA identifies factors covering well known biology of these tumor types. FuncSFA also finds novel factors covering biology whose importance is not yet widely recognized in these tumor types. Some of the factors suggest treatment opportunities that can be further investigated in cell lines and mice.

bioinformatics

Candidate cancer driver mutations in super-enhancers and long-range chromatin interaction networks

A comprehensive catalogue of the mutations that drive tumorigenesis and progression is essential to understanding tumor biology and developing therapies. Protein-coding driver mutations have been well-characterized by large exome-sequencing studies, however many tumors have no mutations in protein-coding driver genes. Non-coding mutations are thought to explain many of these cases, however few non-coding drivers besides TERT promoter are known. To fill this gap, we analyzed 150,000 cis-regulatory regions in 1,844 whole cancer genomes from the ICGC-TCGA PCAWG project. Using our new method, ActiveDriverWGS, we found 41 frequently mutated regulatory elements (FMREs) enriched in non-coding SNVs and indels (FDR<0.05) characterized by aging-associated mutation signatures and frequent structural variants. Most FMREs are distal from genes, reported here for the first time and also recovered by additional driver discovery methods. FMREs were enriched in super-enhancers, H3K27ac enhancer marks of primary tumors and long-range chromatin interactions, suggesting that the mutations drive cancer by distally controlling gene expression through threedimensional genome organization. In support of this hypothesis, the chromatin interaction network of FMREs and target genes revealed associations of mutations and differential gene expression of known and novel cancer genes (e.g., CNNB1IP1, RCC1), activation of immune response pathways and altered enhancer marks. Thus distal genomic regions may include additional, infrequently mutated drivers that act on target genes via chromatin loops. Our study is an important step towards finding such regulatory regions and deciphering the somatic mutation landscape of the non-coding genome.

cancer biology

Novel insights into the molecular heterogeneity of hepatocellularcarcinoma

Hepatocellular carcinoma (HCC) is influenced by numerous factors, which results in diverse genetic, epigenetic and transcriptional scenarios, thus posing obvious challenges for disease management. We scrutinized the molecular heterogeneity of HCC with a multi-omics approach in two small cohorts of resected and explanted livers. Whole-genome transcriptomics was conducted, including polyadenylated transcripts and micro (mi)-RNAs. Copy number variants (CNV) were inferred from whole genome low-pass sequencing data. Fifty-six cancer-related genes were screened using an oncology panel assay. HCC was associated with a dramatic transcriptional deregulation of hundreds of protein-coding genes suggesting downregulation of drugs catabolism, induction of inflammatory responses, and increased cell proliferation in resected livers. Moreover, several long non-coding RNAs and miRNAs not reported previously in the context of HCC were found deregulated. In explanted livers, downregulation of genes involved in energy-producing processes and upregulation of genes aiding in glycolysis were detected. Numerous CNV events were observed, with conspicuous hotspots on chromosomes 1 and 17. Amplifications were more common than deletions, and spanned regions containing genes potentially involved in tumorigenesis. CSF1R, FGFR3, FLT3, NPM1, PDGFRA, PTEN, SMO and TP53 were mutated in all tumors, while other 26 cancer-related genes were mutated with variable penetrance. Our results highlight a remarkable molecular heterogeneity between HCC tumors and reinforce the notion that precision medicine approaches are urgently needed for cancer treatment. We expect that our results will serve as a valuable dataset that will generate hypotheses for us or other researchers to evaluate to ultimately improve our understanding of HCC biology.

cancer biology

Integrated single-nucleotide and structural variation signatures of DNA-repair deficient human cancers

Mutation signatures in cancer genomes reflect endogenous and exogenous mutational processes, offering insights into tumour etiology, features for prognostic and biologic stratification and vulnerabilities to be exploited therapeutically. We present a novel machine learning formalism for improved signature inference, based on multi-modal correlated topic models (MMCTM) which can at once infer signatures from both single nucleotide and structural variation counts derived from cancer genome sequencing data. We exemplify the utility of our approach on two hormone driven, DNA repair deficient cancers: breast and ovary (n=755 cases total). Our results illuminate a new age-associated structural variation signature in breast cancer, and an independently identified substructure within homologous recombination deficient (HRD) tumours in breast and ovarian cancer. Together, our study emphasizes the importance of integrating multiple mutation modes for signature discovery and patient stratification, with biological and clinical implications for DNA repair deficient cancers.

cancer biology

Clusterization in head and neck squamous carcinomas based on lncRNA expression: molecular and clinical correlates

BackgroundLong non-coding RNAs (lncRNAs) have emerged as key players in a remarkably variety of biological processes and pathologic conditions, including cancer. Next-generation sequencing technologies and bioinformatics procedures predict the existence of tens of thousands of lncRNAs, from which we know the functions of only a handful of them, and very little is known in cancer types such as head and neck squamous cell carcinomas (HNSCCs).\n\nResultsHere, we use RNA-seq expression data from The Cancer Genome Atlas (TCGA) and various statistic and software tools in order to get insight about the lncRNome in HNSCC. Based on lncRNAs expression across 426 samples, we discover five distinct tumor clusters that we compare with reported clusters based on various genomic/genetic features. Results demonstrate significant associations between lncRNA-based clustering and DNA-methylation, TP53 mutation, and human papillomavirus infection. Using \"guilt by association\" procedures, we infer the possible biological functions of representative lncRNAs of each cluster. Furthermore, we found that lncRNA clustering is correlated with some important clinical and pathologic features, including patient survival after treatment, tumor grade or sub-anatomical location.\n\nConclusionsWe present a landscape of lncRNAs in HNSCC, and provide associations with important genotypic and phenotypic features that may help to understand the disease.

genomics

An integrative systems biology and experimental approach identifies convergence of epithelial plasticity, metabolism, and autophagy to promote chemoresistance

The evolution of therapeutic resistance is a major cause of death for patients with solid tumors. The development of therapy resistance is shaped by the ecological dynamics within the tumor microenvironment and the selective pressure induced by the host immune system. These ecological and selective forces often lead to evolutionary convergence on one or more pathways or hallmarks that drive progression. These hallmarks are, in turn, intimately linked to each other through gene expression networks. Thus, a deeper understanding of the evolutionary convergences that occur at the gene expression level could reveal vulnerabilities that could be targeted to treat therapy-resistant cancer. To this end, we used a combination of phylogenetic clustering, systems biology analyses, and wet-bench molecular experimentation to identify convergences in gene expression data onto common signaling pathways. We applied these methods to derive new insights about the networks at play during TGF-{beta}-mediated epithelial-mesenchymal transition in a lung cancer model system. Phylogenetics analyses of gene expression data from TGF-{beta} treated cells revealed evolutionary convergence of cells toward amine-metabolic pathways and autophagy during TGF-{beta} treatment. Using high-throughput drug screens, we found that knockdown of the autophagy regulatory, ATG16L1, re-sensitized lung cancer cells to cancer therapies following TGF-{beta}-induced resistance, implicating autophagy as a TGF-{beta}-mediated chemoresistance mechanism. Analysis of publicly-available clinical data sets validated the adverse prognostic importance of ATG16L expression in multiple cancer types including kidney, lung, and colon cancer patients. These analyses reveal the usefulness of combining evolutionary and systems biology methods with experimental validation to illuminate new therapeutic vulnerabilities.

cancer biology

Single-domain antibodies represent novel alternatives to monoclonal antibodies as targeting agents against the human papillomavirus 16 E6 protein

Approximately one-fifth of all malignancies worldwide are etiologically-associated with a persistent viral or bacterial infection. Thus, there is particular interest in therapeutic molecules which utilize components of a natural immune response to specifically inhibit oncogenic microbial proteins, as it is anticipated they will elicit fewer off-target effects than conventional treatments. This concept has been explored in the context of human papillomavirus type 16 (HPV16)-related cancers, through the development of monoclonal antibodies and fragments thereof against the viral E6 oncoprotein. However, challenges related to the biology of E6 as well as the functional properties of the antibodies themselves appear to have precluded their clinical translation. In this study, we attempted to address these issues by exploring the utility of the variable domains of camelid heavy-chain-only antibodies (denoted as VHHs). Through the construction and panning of two llama immune VHH phage display libraries, a pool of potential VHHs was isolated. The interactions of these VHHs with recombinant E6 protein were further characterized using ELISA, Western blotting under both denaturing and native conditions, as well as surface plasmon resonance, and three antibodies were identified that bound recombinant E6 with affinities in the nanomolar range. Our results now lead the way for subsequent studies into the ability of these novel molecules to inhibit HPV16-infected cells in vitro and in vivo.

cancer biology

Cancer subtype identification using somatic mutation data

BACKGROUNDWith the onset of next generation sequencing technologies, we have made great progress in identifying recurrent mutational drivers of cancer. As cancer tissues are now frequently screened for specific sets of mutations, a large amount of samples has become available for analysis. Classification of patients with similar mutation profiles may help identifying subgroups of patients who might benefit from specific types of treatment. However, classification based on somatic mutations is challenging due to the sparseness and heterogeneity of the data.\n\nMETHODSHere, we describe a new method to de-sparsify somatic mutation data using biological pathways. We applied this method to 23 cancer types from The Cancer Genome Atlas, including samples from 5, 805 primary tumors.\n\nRESULTSWe show that, for most cancer types, de-sparsified mutation data associates with phenotypic data. We identify poor prognostic subtypes in three cancer types, which are associated with mutations in signal transduction pathways for which targeted treatment options are available. We identify subtype-drug associations for 14 additional subtypes. Finally, we perform a pan-cancer subtyping analysis and identify nine pan-cancer subtypes, which associate with mutations in four overarching sets of biological pathways.\n\nCONCLUSIONSThis study is an important step towards understanding mutational patterns in cancer.

genomics

Oxygen diffusion in ellipsoidal tumor spheroids

Oxygen plays a central role in cellular metabolism, in both healthy and tumour tissue. The presence and concentration of molecular oxygen in tumours has a substantial effect on both radiotherapy response and tumour evolution, and as a result the oxygen micro-environment is an area of intense research interest. Multicellular tumour spheroids closely mimic real avascular tumours, and in particular they exhibit physiologically relevant heterogeneous oxygen distribution. This property has made them a vital part of in vitro experimentation. For ideal spheroids, their heterogeneous oxygen distributions can be predicted from theory, allowing determination of cellular oxygen consumption rate (OCR) and anoxic extent. However, experimental tumour spheroids often depart markedly from perfect sphericity. There has been little consideration of this reality. To date, the question of how far an ellipsoid can diverge from perfect sphericity before spherical assumptions breakdown remains unanswered. In this work we derive equations governing oxygen distribution (and more generally, nutrient and drug distribution) in both prolate and oblate tumour ellipsoids, and quantify the theoretical limits of the assumption that the spheroid is a perfect sphere. Results of this analysis yield new methods for quantifying OCR in ellipsoidal spheroids, and how this can be applied to markedly increase experimental throughput and quality.\n\nAuthor summaryMulticellular tumour spheroids (MCTS) are an increasingly important tool in cancer research, exhibiting non-homogeneous oxygen distributions and central necrosis. These are more similar to in situ avascular tumours than conventional 2D biology, rendering them exceptionally useful experimental models. Analysis of spheroids can yield vital information about cellular oxygen consumption rates, and the heterogeneous oxygen contribution. However, such analysis pivots on the assumption of perfect sphericity, when in reality spheroids often depart from such an ideal. In this work, we construct a theoretical oxygen diffusion model for ellipsoidal tumour spheroids in both prolate and oblate geometries. With these models established, we quantify the limits of the spherical assumption, and illustrate the effect of this assumption breaking down. Methods of circumventing this breakdown are also presented, and the analysis here suggests new methods for expanding experimental throughput to also include ellipsoidal data.

cancer biology

Driver Pattern Identification Over The Gene Co-Expression Of Drug Response In Ovarian Cancer By Integrating High Throughput Genomics Data

The multiple types of high throughput genomics data create a potential opportunity to identify driver pattern in ovarian cancer, which will acquire some novel and clinical biomarkers for appropriate diagnosis and treatment to cancer patients. However, it is a great challenging work to integrate omics data, including somatic mutations, Copy Number Variations (CNVs) and gene expression profiles, to distinguish interactions and regulations which are hidden in drug response dataset of ovarian cancer. To distinguish the candidate driver genes and the corresponding driving pattern for resistant and sensitive tumor from the heterogeneous data, we combined gene co-expression modules and mutation modulators and proposed the identification driver patterns method. Firstly, co-expression network analysis is applied to explore gene modules for gene expression profiles via weighted correlation network analysis (WGCNA). Secondly, mutation matrix is generated by integrating the CNVs and somatic mutations, and a mutation network is constructed from this mutation matrix. The candidate modulators are selected from the significant genes by clustering the vertex of the mutation network. At last, regression tree model is utilized for module networks learning in which the achieved gene modules and candidate modulators are trained for the driving pattern identification and modulator regulatory exploring. Many of the candidate modulators identified are known to be involved in biological meaningful processes associated with ovarian cancer, which can be regard as potential driver genes, such as CCL11, CCL16, CCL18, CCL23, CCL8, CCL5, APOB, BRCA1, SLC18A1, FGF22, GADD45B, GNA15, GNA11 and so on, which can help to facilitate the discovery of biomarkers, molecular diagnostics, and drug discovery.

bioinformatics

Identification of an immune gene expression signature associated with favorable clinical features in Treg-enriched patient tumor samples

Immune heterogeneity within the tumor microenvironment undoubtedly adds several layers of complexity to our understanding of drug sensitivity and patient prognosis across various cancer types. Within the tumor microenvironment, immunogenicity is a favorable clinical feature in part driven by the antitumor activity of CD8+ T cells. However, tumors often inhibit this antitumor activity by exploiting the suppressive function of Regulatory T cells (Tregs), thus suppressing the adaptive immune response. Despite the seemingly intuitive immunosuppressive biology of Tregs, prognostic studies have produced contradictory results regarding the relationship between Treg enrichment and survival. We therefore analyzed RNA-seq data of Treg-enriched tumor samples to derive a pan-cancer gene signature able to help reconcile the inconsistent results of Treg studies, by better understanding the variable clinical association of Tregs across alternative tumor contexts. We show that increased expression of a 32-gene signature in Treg-enriched tumor samples (n=135) is able to distinguish a cohort of patients associated with chemosensitivity and overall survival This cohort is also enriched for CD8+ T cell abundance, as well as the antitumor M1 macrophage subtype. With a subsequent validation in a larger TCGA pool of Treg-enriched patients (n = 626), our results reveal a gene signature able to produce unsupervised clusters of Treg-enriched patients, with one cluster of patients uniquely representative of an immunogenic tumor microenvironment. Ultimately, these results support the proposed gene signature as a putative biomarker to identify certain Treg-enriched patients with immunogenic tumors that are more likely to be associated with features of favorable clinical outcome.

cancer biology

Intraductal patient derived xenografts of estrogen receptor positive breast cancer recapitulate the histopathological spectrum and metastatic potential of human lesions

Estrogen receptor positive (ER+) or \"luminal\" breast cancers were notoriously difficult to establish as patient-derived xenografts (PDXs). We and others recently demonstrated that the microenvironment is critical for ER+ tumor cells; by grafting them into milk ducts >90% take rates are achieved and many features of the human disease are recapitulated. This intra-ductal (ID) approach holds promise for personalized medicine, yet human and murine stroma are organized differently and this and other species specificities may limit the value of this model. Here, we analyzed 21 ER+ ID-PDXs histopathologically. We find that ID-PDXs vary in extent and define four histopathological patterns: flat, lobular, in situ, and invasive, which occur in pure and combined forms. The ID-PDXs replicate earlier stages of tumor development than their clinical counterparts. Micrometastases are already detected when lesions appear in situ. Tumor extent, histopathological patterns, and metastatic load correlate with biological properties of their tumors of origin. Our findings add evidence to the validity of the intraductal model for in vivo studies of ER+ breast cancer and raise the intriguing possibility that tumor cell dissemination may occur earlier than currently thought.\n\nConflict of interest statementThe authors declare no conflict of interest.

cancer biology

Unique genomic features and deeply-conserved functions of long non-coding RNAs in the Cancer LncRNA Census (CLC)

Long non-coding RNAs (lncRNAs) that drive tumorigenesis are a growing focus of cancer genomics studies. To facilitate further discovery, we have created the \"Cancer LncRNA Census\" (CLC), a manually-curated and strictly-defined compilation of lncRNAs with causative roles in cancer. CLC has two principle applications: first, as a resource for training and benchmarking de novo identification methods; and second, as a dataset for studying the fundamental properties of these genes.\n\nCLC Version 1 comprises 122 lncRNAs implicated in 29 distinct cancers. LncRNAs are included based on functional or genetic evidence for causative roles in cancer progression. All belong to the GENCODE reference annotation, to enable integration across projects and datasets. For each entry, the evidence type, biological activity (oncogene or tumour suppressor), source reference and cancer type are recorded. Supporting its usefulness, CLC genes are significantly enriched amongst de novo predicted driver genes from PCAWG. CLC genes are distinguished from other lncRNAs by a series of features consistent with biological function, including gene length, high expression and sequence conservation of both exons and promoters. We identify a trend for CLC genes to be co-localised with known protein-coding cancer genes along the human genome. Finally, by integrating data from transposon-mutagenesis functional screens, we show that mouse orthologues of CLC genes tend also to be cancer genes.\n\nThus CLC represents a valuable resource for research into long non-coding RNAs in cancer. Their evolutionary and genomic properties have implications for understanding disease mechanisms and point to conserved functions across ~80 million years of evolution.

bioinformatics

LOTUS: a Single- and Multitask Machine Learning Algorithm for the Prediction of Cancer Driver Genes

Cancer driver genes, i.e., oncogenes and tumor suppressor genes, are involved in the acquisition of important functions in tumors, providing a selective growth advantage, allowing uncontrolled proliferation and avoiding apoptosis. It is therefore important to identify these driver genes, both for the fundamental understanding of cancer and to help finding new therapeutic targets. Although the most frequently mutated driver genes have been identified, it is believed that many more remain to be discovered, particularly for driver genes specific to some cancer types.\n\nIn this paper we propose a new computational method called LOTUS to predict new driver genes. LOTUS is a machine-learning based approach which allows to integrate various types of data in a versatile manner, including informations about gene mutations and protein-protein interactions. In addition, LOTUS can predict cancer driver genes in a pan-cancer setting as well as for specific cancer types, using a multitask learning strategy to share information across cancer types.\n\nWe empirically show that LOTUS outperforms three other state-of-the-art driver gene prediction methods, both in terms of intrinsic consistency and prediction accuracy, and provide predictions of new cancer genes across many cancer types.\n\nAuthor summaryCancer development is driven by mutations and dysfunction of important, so-called cancer driver genes, that could be targeted by targeted therapies. While a number of such cancer genes have already been identified, it is believed that many more remain to be discovered. To help prioritize experimental investigations of candidate genes, several computational methods have been proposed to rank promising candidates based on their mutations in large cohorts of cancer cases, or on their interactions with known driver genes in biological networks. We propose LOTUS, a new computational approach to identify genes with high oncogenic potential. LOTUS implements a machine learning approach to learn an oncogenic potential score from known driver genes, and brings two novelties compared to existing methods. First, it allows to easily combine heterogeneous informations into the scoring function, which we illustrate by learning a scoring function from both known mutations in large cancer cohorts and interactions in biological networks. Second, using a multitask learning strategy, it can predict different driver genes for different cancer types, while sharing information between them to improve the prediction for every type. We provide experimental results showing that LOTUS significantly outperforms several state-of-the-art cancer gene prediction softwares.

bioinformatics