Search bioRxivSearch

Biology subjects

Haibe-Kains, B.

Publications and source records attributed to Haibe-Kains, B..

14 recordsLinked to original sources

Modeling cellular response in large-scale radiogenomic databases to advance precision radiotherapy

Radiotherapy is integral to the care of a majority of cancer patients. Despite differences in tumor responses to radiation (radioresponse), dose prescriptions are not currently tailored to individual patients. Recent large-scale cancer cell line databases hold the promise of unravelling the complex molecular arrangements underlying cellular response to radiation, which is critical to novel predictive biomarker discovery. Here, we present RadioGx, a computational platform for integrative analyses of radioresponse using radiogenomic databases. We first used RadioGx to investigate the robustness of radioresponse assays and indicators. We then combined radioresponse and genome-wide molecular data with established radiobiological models to predict molecular pathways that are relevant for individual tissue types and conditions. We also applied RadioGx to pharmacogenomic data to identify several classes of drugs whose effects correlate with radioresponse. RadioGx provides a unique computational toolbox to advance preclinical research for radiation oncology and precision medicine.

bioinformatics

Chromatin Blueprint Of Glioblastoma Stem Cells Reveals Common Drug Candidates For Distinct Subtypes

Chromatin accessibility discriminates stem from mature cell populations, enabling the identification of primitive stem-like cells in primary tumors, such as Glioblastoma (GBM) where self-renewing cells driving cancer progression and recurrence are prime targets for therapeutic intervention. We show, using single-cell chromatin accessibility, that primary GBMs harbor a heterogeneous self-renewing population whose diversity is captured in patient-derived glioblastoma stem cells (GSCs). In depth characterization of chromatin accessibility in GSCs identifies three GSC states: Reactive, Constructive, and Invasive, each governed by uniquely essential transcription factors and present within GBMs in varying proportions. Orthotopic xenografts reveal that GSC states associate with survival, and identify an invasive GSC signature predictive of low patient survival. Our chromatin-driven characterization of GSC states improves prognostic precision and identifies dependencies to guide combination therapies.

cancer biology

Meta-analysis of 1,200 transcriptomic profiles identifies a prognostic model for pancreatic ductal adenocarcinoma

BackgroundWith a dismal 8% median 5-year overall survival (OS), pancreatic ductal adenocarcinoma (PDAC) is highly lethal. Only 10-20% of patients are eligible for surgery, and over 50% of these will die within a year of surgery. Identify molecular predictors of early death would enable the selection of PDAC patients at high risk.\n\nMethodsWe developed the Pancreatic Cancer Overall Survival Predictor (PCOSP), a prognostic model built from a unique set of 89 PDAC tumors where gene expression was profiled using both microarray and sequencing platforms. We used a meta-analysis framework based on the binary gene pair method to create gene expression barcodes robust to biases arising from heterogeneous profiling platforms and batch effects. Leveraging the largest compendium of PDAC transcriptomic datasets to date, we show that PCOSP is a robust single-sample predictor of early death ([≤]1 yr) after surgery in a subset of 823 samples with available transcriptomics and survival data.\n\nResultsThe PCOSP model was strongly and significantly prognostic with a meta-estimate of the area under the receiver operating curve (AUROC) of 0.70 (P=1.9e-18) and hazard ratio (HR) of 1.95(1.6-2.3) (P=2.6e-16) for binary and survival predictions, respectively. The prognostic value of PCOSP was independent of clinicopathological parameters and molecular subtypes. Over-representation analysis of the PCOSP 2619 gene-pairs (1070 unique genes) unveiled pathways associated with Hedgehog signalling, epithelial mesenchymal transition (EMT) and extracellular matrix (ECM) signalling.\n\nConclusionPCOSP could improve treatment decision by identifying patients who will not benefit from standard surgery/chemotherapy and may benefit from alternate approaches.\n\nAbbreviations

genomics

Metabolic adaptations underlie epigenetic vulnerabilities in chemoresistant breast cancer.

Cancer cell survival upon cytotoxic drug exposure leads to changes in cell identity, dictated by the epigenome. Several metabolites serve as substrates or co-factors to chromatin-modifying enzymes, suggesting that metabolic changes can underlie change in cell fate. Here, we show that progression of triple-negative breast cancer (TNBC) to taxane-resistance is characterized by altered methionine metabolism and S-adenosylmethionine (SAM) availability, giving rise to DNA hypomethylation in regions enriched for transposable elements (TE). Compensatory redistribution of H3K27me3 forming Large Organized Chromatin domains of lysine (K) modification (LOCK) prevents expression of TE in taxane-resistant cells. Pharmacological inhibition of EZH2, the H3K27me3 methyltransferase, alleviates TE repression, leading to the accumulation of dsRNA and activation of the interferon viral mimicry-response, specifically inhibiting the growth of taxane-resistant TNBC. Together, our work delineates a role for metabolic adaptations in redefining the epigenome of taxane-resistant TNBC cells and underlies an epigenetic vulnerability toward pharmacological inhibition of EZH2.

genomics

Genome analysis and data sharing informs timing of molecular events in pancreatic neuroendocrine tumour

Neuroendocrine tumours (NETs) are rare, slow growing cancers that present in a diversity of tissues. To understand molecular underpinnings of gastrointestinal (GINET) and pancreatic NETs (PNETs), we profiled 45 tumours combining exome, RNA, and shallow whole genome sequencing, as well as fluorescent in situ hybridization. In addition to expected somatic mutations and copy number alterations, we found that PNETs contained a highly consistent copy neutral loss-of-heterozygosity (CN-LOH) profile affecting over half of the genome; a greater percentage than any cancer analyzed to date. Our data indicates that onset of extreme autozygosity may be progressive, associated with metastasis, and initially triggered by the loss of DAXX/ATRX, and subsequent biallelic loss of MEN1. We confirmed this molecular timing model using targeted clinical sequencing data from an additional 43 NETs made available by the AACR GENIE project. Against this background of CN-LOH, several chromosomal regions consistently retained heterozygosity, suggesting selection for crucial allele-specific components specific to PNET progression and potential new therapeutic targets.\n\nStatement of significanceWe have discovered that pancreatic neuroendocrine tumours contain a characteristic pattern of copy neutral loss-of-heterozygosity affecting the majority of the genome following mutations of MEN1 and ATRX/DAXX. Against this background of loss-of-heterozygosity, specific genomic regions are consistently retained and may therefore contain vulnerable therapeutic targets for pancreatic neuroendocrine tumours.

cancer biology

CREAM: Clustering of genomic REgions Analysis Method

Cellular identity relies on cell type-specific gene expression profiles controlled by cis-regulatory elements (CREs), such as promoters, enhancers and anchors of chromatin interactions. CREs are unevenly distributed across the genome, giving rise to distinct subsets such as individual CREs and Clusters Of cis-Regulatory Elements (COREs), also known as super-enhancers. Identifying COREs is a challenge due to technical and biological features that entail variability in the distribution of distances between CREs within a given dataset. To address this issue, we developed a new unsupervised machine learning approach termed Clustering of genomic REgions Analysis Method (CREAM) that outperforms the Ranking Of Super Enhancer (ROSE) approach. Specifically CREAM identified COREs are enriched in CREs strongly bound by master transcription factors according to ChIP-seq signal intensity, are proximal to highly expressed genes, are preferentially found near genes essential for cell growth and are more predictive of cell identity. Moreover, we show that CREAM enables subtyping primary prostate tumor samples according to their CORE distribution across the genome. We further show that COREs are enriched compared to individual CREs at TAD boundaries and these are preferentially bound by CTCF and factors of the cohesin complex (e.g.: RAD21 and SMC3). Finally, using CREAM against transcription factor ChIP-seq reveals CTCF and cohesin-specific COREs preferentially at TAD boundaries compared to intra-TADs. CREAM is available as an open source R package (https://CRAN.R-project.org/package=CREAM) to identify COREs from cis-regulatory annotation datasets from any biological samples.

bioinformatics

Whole Genomes Define Concordance of Matched Primary, Xenograft, and Organoid Models of Pancreas Cancer

Pancreatic ductal adenocarcinoma (PDAC) has the worst prognosis among solid malignancies and improved therapeutic strategies are needed to improve outcomes. Patient-derived xenografts (PDX) and patient-derived organoids (PDO) serve as promising tools to identify new drugs with therapeutic potential in PDAC. For these preclinical disease models to be effective, they should both recapitulate the molecular heterogeneity of PDAC and validate patient-specific therapeutic sensitivities. To date however, deep characterization of PDAC PDX and PDO models and comparison with matched human tumour remains largely unaddressed at the whole genome level. We conducted a comprehensive assessment of the genetic landscape of 16 whole-genome pairs of tumours and matched PDX, from primary PDAC and liver metastasis, including a unique cohort of 5 trios of matched primary tumour, PDX, and PDO. We developed a new pipeline to score concordance between PDAC models and their paired human tumours for genomic events, including mutations, structural variations, and copy number variations. Comparison of genomic events in the tumours and matched disease models displayed single-gene concordance across major PDAC driver genes, and genome-wide similarities of copy number changes. Genome-wide and chromosome-centric analysis of structural variation (SV) events revealed high variability across tumours and disease models, but also highlighted previously unrecognized concordance across chromosomes that demonstrate clustered SV events. Our approach and results demonstrate that PDX and PDO recapitulate PDAC tumourigenesis with respect to simple somatic mutations and copy number changes, and capture major SV events that are found in both resected and metastatic tumours.

bioinformatics

PharmacoDB: an integrative database for mining in vitro anticancer drug screening studies

Recent pharmacogenomic studies profiled large panels of cancer cell lines against hundreds of approved drugs and experimental chemical compounds. The overarching goal of these screens is to measure sensitivity of cell lines to chemical perturbation, correlate these measures to genomic features, and thereby develop novel predictors of drug response. However, leveraging this valuable data is challenging due to the lack of standards for annotating cell lines and chemical compounds, and quantifying drug response. Moreover, it has been recently shown that the complexity and complementarity of the experimental protocols used in the field result in high levels of technical and biological variation in the in vitro pharmacological profiles. There is therefore a need for new tools to facilitate rigorous comparison and integrative analysis of large-scale drug screening datasets. To address this issue, we have developed PharmacoDB (pharmacodb.pmgenomics.ca), a database integrating the largest pharmacogenomic studies published to date. Here, we describe how the curation of cell line and chemical compound identifiers maximizes the overlap between datasets and how users can leverage such data to compare and extract robust drug phenotypes. PharmacoDB provides a unique resource to mine a compendium of curated pharmacogenomic datasets that are otherwise disparate and difficult to integrate.\n\nKey pointsO_LICuration of cell line and drug identifiers in the largest pharmacogenomic studies published to date\nC_LIO_LIUniform processing of drug sensitivity data to reduce heterogeneity across studies\nC_LIO_LIMultiple drug response summary metrics enabling visual comparison and integrative analysis\nC_LI

bioinformatics

Consensus on Molecular Subtypes of Ovarian Cancer

INTRODUCTIONVarious computational methods for gene expression-based subtyping of high-grade serous (HGS) ovarian cancer have been proposed. This resulted in the identification of molecular subtypes that are based on different datasets and were differentially validated, making it difficult to achieve consensus on which definitions to use in follow-up studies. We assess three major subtype classifiers for their robustness and association to outcome by a meta-analysis of publicly available expression data, and provide a classifier that represents their consensus.\n\nMETHODSWe use a compendium of 15 microarray datasets consisting of 1,774 HGS ovarian tumors to assess 1) concordance between published subtyping algorithms, 2) robustness of those algorithms to re-clustering across datasets, and 3) association of subtypes with overall survival. A consensus classifier is trained on concordantly classified samples, and validated by leave-one-dataset-out validation.\n\nRESULTSEach subtyping classifier identified subsets significantly differing in overall survival, but were not robust to re-fitting in independent datasets and grouped only approximately one third of patients concordantly into four subtypes. We propose a consensus classifier to identify the minority of unambiguously classifiable tumors across multiple gene expression platforms, using a 100-gene signature. The resulting consensus subtypes correlate with patient age, survival, tumor purity, and lymphocyte infiltration.\n\nCONCLUSIONSOur analysis demonstrates that most HGS ovarian cancers are not able to be subtyped. A minority of tumors can be classified and our proposed consensus classifier consolidates and improves on the robustness of three previously proposed subtype classifiers. It provides reliable stratification of patients with HGS ovarian tumors of clearly defined subtype, and will assist in studying the role of polyclonality in the majority of tumors that are not robustly classifiable.

cancer biology

Gene isoforms as expression-based biomarkers predictive of drug response in vitro

BackgroundOne of the main challenges in precision medicine is the identification of molecular features associated to drug response to provide clinicians with tools to select the best therapy for each individual cancer patient. The recent adoption of next-generation sequencing technologies enables accurate profiling of not only gene expression but also alternatively-spliced transcripts in large-scale pharmacogenomic studies. Given that altered mRNA splicing has been shown to be prominent in cancers, linking this feature to drug response will open new avenues of research in biomarker discovery.\n\nMethodsTo address the lack of reproducibility of drug sensitivity measurements across studies, we developed a meta-analytical framework combining the pharmacological data generated within the Cancer Cell Line Encyclopedia (CCLE) and the Genomics of Drug Sensitivity in Cancer (GDSC). Predictive models are fitted with CCLE RNA-seq data as predictor variables, controlled for tissue type, and combined GDSC and CCLE drug sensitivity values as dependent variables.\n\nResultsWe first validated the biomarkers identified from GDSC and CCLE using an existing pharmacogenomic dataset of 70 breast cancer cell lines. We further selected four drugs with the most promising biomarkers to test whether their predictive value is robust to change in pharmacological assay. We successfully validated 10 isoform-based biomarkers predictive of drug response in breast cancer, including TGFA-001 for the MEK tyrosine kinase inhibitor (TKI) AZD6244, DUOX-001 for the EGFR inhibitor erlotinib, and CPEB4-001 transcript expression associated with lack of sensitivity to paclitaxel.\n\nConclusionThe results of our meta-analysis of pharmacogenomic data suggest that isoforms represent a rich resource for biomarkers predictive of response to chemo- and targeted therapies. Our study also showed that the validation rate for this type of biomarkers is low (<50%) for most drugs, supporting the requirements for independent datasets to identify reproducible predictors of response to anticancer drugs.

bioinformatics

Software For The Integration Of Multi-Omics Experiments In Bioconductor

Multi-omics experiments are increasingly commonplace in biomedical research, and add layers of complexity to experimental design, data integration, and analysis. R and Bioconductor provide a generic framework for statistical analysis and visualization, as well as specialized data classes for a variety of high-throughput data types, but methods are lacking for integrative analysis of multi-omics experiments. The MultiAssayExperiment software package, implemented in R and leveraging Bioconductor software and design principles, provides for the coordinated representation of, storage of, and operation on multiple diverse genomics data. We provide all of the multiple omics data for each cancer tissue in The Cancer Genome Atlas (TCGA) as ready-to-analyze MultiAssayExperiment objects, and demonstrate in these and other datasets how the software simplifies data representation, statistical analysis, and visualization. The MultiAssayExperiment Bioconductor package reduces major obstacles to efficient, scalable and reproducible statistical analysis of multi-omics data and enhances data science applications of multiple omics datasets.

bioinformatics

Similarity identification in gene expression patterns as a new approach in phenotype classification

Stratifying healthy and malignant phenotypes and identifying their biological states using high-throughput molecular data has been the focus of many computational approaches during the last decade. Using multivariate changes in expression of genes within biological pathways, as fingerprints of complex phenotypes, we developed a new methodology for Similarity Identification in Gene expressioN (SIGN). In this approach, we use centroid classifier to identify phenotype of each biological sample. To obtain similarity of a given biological sample with classes of phenotypes, we defined a new distance measure, transcriptional similarity coefficient (TSC) which captures similarity of gene expression patterns between a biological pathway in two samples or populations. We showed that TSC, as an interpretable and stable distance measure in SIGN, captures all oncogenic hallmarks for breast cancer even with low sample size, by comparing healthy and patient tumor samples in the largest breast cancer dataset. In this study, we demonstrate that SIGN is a flexible, yet robust approach for classification based on transcriptomics data. Comparing early and late relapses within each molecular subtypes of breast cancer, our method enabled subtype-specific stratification of breast cancer patients into groups with significantly different survival. Moreover, we used SIGN to classify with more than 99% specificity the site of extraction of healthy and tumor samples from the Genotype-Tissue Expression (GTEx) and The Cancer Genome Atlas (TCGA) datasets. We showed that SIGN also enables robust identification of hematopoietic stem cell and progenitors within the hematopoietic hierarchy. We further explored chemical perturbation data in the Connectivity Map (CMAP) database and showed that SIGN was able to classify seven classes of drugs based on their mechanism of action. In conclusion, we showed that SIGN can be used to achieve interpretable and robust transcriptomic-based classification of healthy and malignant samples, as well as drugs based on their known mechanism of action, supporting the generalizability and relevance of the method for the analysis of gene expression profiles.

bioinformatics

CrosstalkNet: mining large-scale bipartite co-expression networks to characterize epi-stroma crosstalk

BackgroundOver the last several years, we have witnessed the metamorphosis of network biology from being a mere representation of molecular interactions to models enabling inference of complex biological processes. Networks provide promising tools to elucidate intercellular interactions that contribute to the functioning of key biological pathways in a cell. However, the exploration of these large-scale networks remains a challenge due to their high-dimensionality.\n\nResultsCrosstalkNet is a user friendly, web-based network visualization tool to retrieve and mine interactions in large-scale bipartite co-expression networks. In this study, we discuss the use of gene co-expression networks to explore the rewiring of interactions between tumor epithelial and stromal cells. We show how CrosstalkNet can be used to efficiently visualize, mine, and interpret large co-expression networks representing the crosstalk occurring between the tumour and its microenvironment.\n\nConclusionCrosstalkNet serves as a tool to assist biologists and clinicians in exploring complex, large interaction graphs to obtain insights into the biological processes that govern the tumor epithelial-stromal crosstalk. A comprehensive tutorial along with case studies are provided with the application.\n\nAvailabilityThe web-based application is available at the following location: http://epistroma.pmgenomics.ca/app/. The code is open-source and freely available from http://github.com/bhklab/EpiStroma-webapp.\n\nContactbhaibeka@uhnresearch.ca

systems biology

Tissue specificity of in vitro drug sensitivity

Research in oncology traditionally focuses on specific tissue type from which the cancer develops. However, advances in high-throughput molecular profiling technologies have enabled the comprehensive characterization of molecular aberrations in multiple cancer types. It was hoped that these large-scale datasets would provide the foundation for a paradigm shift in oncology which would see tumors being classified by their molecular profiles rather than tissue types, but tumors with similar genomic aberrations may respond differently to targeted therapies depending on their tissue of origin. There is therefore a need to reassess the potential association between pharmacological response and tissue of origin for therapeutic drugs, and to test how these associations translate from preclinical to clinical settings.\n\nIn this paper, we investigate the tissue specificity of drug sensitivities in large-scale pharmacological studies and compare these associations to those found in clinical trial descriptions. Our meta-analysis of the four largest in vitro drug screening datasets indicates that tissue of origin is strongly associated with drug response. We identify novel tissue-drug associations, which may present exciting new avenues for drug repurposing. One caveat is that the vast majority of the significant associations found in preclinical settings do not concur with clinical observations. Accordingly, our results call for more testing to find the root cause of the discrepancies between preclinical and clinical observations.

bioinformatics