Search bioRxivSearch

Biology subjects

Collins, C. C.

Publications and source records attributed to Collins, C. C..

4 recordsLinked to original sources

Combinatorial Detection of Conserved Alteration Patterns for Identifying Cancer Subnetworks

BackgroundAdvances in large scale tumor sequencing have lead to an understanding that there are combinations of genomic and transcriptomic alterations speciflc to tumor types, shared across many patients. Unfortunately, computational identiflcation of functionally meaningful shared alteration patterns, impacting gene/protein interaction subnetworks, has proven to be challenging. FindingsWe introduce a novel combinatorial method, cd-CAP, for simultaneous detection of connected subnetworks of an interaction network where genes exhibit conserved alteration patterns across tumor samples. Our method differentiates distinct alteration types associated with each gene (rather than relying on binary information of a gene being altered or not), and simultaneously detects multiple alteration proflle conserved subnetworks. ConclusionsIn a number of The Cancer Genome Atlas (TCGA) data sets, cd-CAP identifled large biologically signiflcant subnetworks with conserved alteration patterns, shared across many tumor samples.

systems biology

BAP1 Loss Predicts Therapeutic Vulnerability in Malignant Peritoneal Mesothelioma

BackgroundMalignant Peritoneal Mesothelioma (PeM) is a rare but frequently fatal cancer that originates from the peritoneal lining of the abdomen. Standard treatment of PeM is limited to cytoreductive surgery and/or chemotherapy, and no effective targeted therapies for PeM yet exist. In the search for novel therapeutic target candidates in PeM, we performed a comprehensive integrative multi-omics analysis of 19 treatment-naive PeM tumors.\n\nResultsThe analysis identified PeM tumors with BAP1 loss to form a distinct molecular subtype characterized by distinct expression patterns of genes involved in chromatin remodeling, DNA repair pathway, and immune checkpoint receptor activation. This PeM subtype could potentially benefit from immune checkpoint, PARP, or HDAC inhibition therapies.\n\nConclusionsOur findings uncover BAP1 as a trackable prognostic and predictive biomarker, and refine PeM disease classification. This integrated molecular characterization provides a comprehensive foundation for developing PeM precision medicine.

cancer biology

Deep Genomic Signature for early metastasis prediction in prostate cancer

For prostate cancer patients, timing and intensity of therapy are adjusted based on their prognosis. Clinical and pathological factors, and recently, gene expression-based signatures have been shown to predict metastatic prostate cancer. Previous studies used labelled datasets, i.e. those with information on the metastasis outcome, to discover gene signatures to predict metastasis. Due to steady progression of prostate cancer, datasets for this cancer have a limited number of labelled samples but more unlabelled samples. In addition to this issue, the high dimensionality of the gene expression data also poses a significant challenge to train a classifier and predict metastasis accurately. In this study, we aim to boost the prediction accuracy by utilizing both labelled and unlabelled datasets together. We propose Deep Genomic Signature (DGS), a method based on Denoising Auto-Encoders (DAEs) and transfer learning. DGS has the following steps: first, we train a DAE on a large unlabelled gene expression dataset to extract the most salient features of its samples. Then, we train another DAE on a small labelled dataset for a similar purpose. Since the labelled dataset is small, we employ a transfer learning approach and use the parameters learned from the first DAE in the second one. This approach enables us to train a large DAE on a small dataset. After training the second DAE, we obtain the list of genes with high weights by applying a standard deviation filter on the transferred and learned weights. Finally, we train an elastic net logistic regression model on the expression of the selected genes to predict metastasis. Because of the elastic net regularization, some of the selected genes have non-zero coefficients in the classifier which we consider as the DGS gene signature for metastasis. We apply DGS to six labelled and one large unlabelled prostate cancer datasets. Results on five validation datasets indicate that DGS outperforms state-of-the-art gene signatures (obtained from only labelled datasets) in terms of prediction accuracy. Survival analyses demonstrate the potential clinical utility of our gene signature that adds novel prognostic information to the well-established clinical factors and the state-of-the-art gene signatures. Finally, pathway analysis reveals that the DGS gene signature captures the hallmarks of prostate cancer metastasis. These results suggest that our method helps to identify a robust gene signature that may improve patient management.

bioinformatics

Computational proteogenomic identification and functional interpretation of translated fusions and micro structural variations in cancer

MotivationRapid advancement in high throughput genome and transcriptome sequencing (HTS) and mass spectrometry (MS) technologies has enabled the acquisition of the genomic, transcriptomic and proteomic data from the same tissue sample. In this paper we introduce a novel computational framework which can integratively analyze all three types of omics data to obtain a complete molecular profile of a tissue sample, in normal and disease conditions. Our framework includes MiStrVar, an algorithmic method we developed to identify micro structural variants (microSVs) on genomic HTS data. Coupled with deFuse, a popular gene fusion detection method we developed earlier, MiStrVar can provide an accurate profile of structurally aberrant transcripts in cancer samples. Given the breakpoints obtained by MiStrVar and deFuse, our framework can then identify all relevant peptides that span the breakpoint junctions and match them with unique proteomic signatures in the respective proteomics data sets. Our framework's ability to observe structural aberrations at three levels of omics data provides means of validating their presence.\n\nResultsWe have applied our framework to all The Cancer Genome Atlas (TCGA) breast cancer Whole Genome Sequencing (WGS) and/or RNA-Seq data sets, spanning all four major subtypes, for which proteomics data from Clinical Proteomic Tumor Analysis Consortium (CPTAC) have been released. A recent study on this dataset focusing on SNVs has reported many that lead to novel peptides [1]. Complementing and significantly broadening this study, we detected 244 novel peptides from 432 candidate genomic or transcriptomic sequence aberrations. Many of the fusions and microSVs we discovered have not been reported in the literature. Interestingly, the vast majority of these translated aberrations (in particular, fusions) were private, demonstrating the extensive inter-genomic heterogeneity present in breast cancer. Many of these aberrations also have matching out-of-frame downstream peptides, potentially indicating novel protein sequence and structure. Moreover, the most significantly enriched genes involved in translated fusions are cancer-related. Furthermore a number of the somatic, translated microSVs are observed in tumor suppressor genes.\n\nContactcenksahi@indiana.edu

bioinformatics