Search bioRxivSearch

Biology subjects

Stein, L. D.

Publications and source records attributed to Stein, L. D..

6 recordsLinked to original sources

Candidate cancer driver mutations in super-enhancers and long-range chromatin interaction networks

A comprehensive catalogue of the mutations that drive tumorigenesis and progression is essential to understanding tumor biology and developing therapies. Protein-coding driver mutations have been well-characterized by large exome-sequencing studies, however many tumors have no mutations in protein-coding driver genes. Non-coding mutations are thought to explain many of these cases, however few non-coding drivers besides TERT promoter are known. To fill this gap, we analyzed 150,000 cis-regulatory regions in 1,844 whole cancer genomes from the ICGC-TCGA PCAWG project. Using our new method, ActiveDriverWGS, we found 41 frequently mutated regulatory elements (FMREs) enriched in non-coding SNVs and indels (FDR<0.05) characterized by aging-associated mutation signatures and frequent structural variants. Most FMREs are distal from genes, reported here for the first time and also recovered by additional driver discovery methods. FMREs were enriched in super-enhancers, H3K27ac enhancer marks of primary tumors and long-range chromatin interactions, suggesting that the mutations drive cancer by distally controlling gene expression through threedimensional genome organization. In support of this hypothesis, the chromatin interaction network of FMREs and target genes revealed associations of mutations and differential gene expression of known and novel cancer genes (e.g., CNNB1IP1, RCC1), activation of immune response pathways and altered enhancer marks. Thus distal genomic regions may include additional, infrequently mutated drivers that act on target genes via chromatin loops. Our study is an important step towards finding such regulatory regions and deciphering the somatic mutation landscape of the non-coding genome.

cancer biology

DriverPower: Combined burden and functional impact tests for cancer driver discovery

We describe DriverPower, a software package that uses mutational burden and functional impact evidence to identify cancer driver mutations in coding and non-coding sites within cancer whole genomes. Using a total of 1,373 genomic features derived from public sources, DriverPowers background mutation model explains up to 93% of the regional variance in the mutation rate across a variety of tumour types. By incorporating functional impact scores, we are able to further increase the accuracy of driver discovery. Testing across a collection of 2,583 cancer genomes from the Pan-Cancer Analysis of Whole Genomes (PCAWG) project, DriverPower identifies 217 coding and 95 non-coding driver candidates. Comparing to six published methods used by the PCAWG Drivers and Functional Interpretation Group, DriverPower has the highest F1-score for both coding and non-coding driver discovery. This demonstrates that DriverPower is an effective framework for computational driver discovery.

bioinformatics

Accurate Discrimination of 23 Major Cancer Types via Whole Genome Somatic Mutation Patterns

The two strongest factors predicting a human cancers clinical behaviour are the primary tumours anatomic organ of origin and its histopathology. However, roughly 3% of the time a cancer presents with metastatic disease and no primary can be determined even after a thorough radiological survey. A related dilemma arises when a radiologically defined mass is sampled by cytology yielding cancerous cells, but the cytologist cannot distinguish between a primary tumour and a metastasis from elsewhere.\n\nHere we use whole genome sequencing (WGS) data from the ICGC/TCGA PanCancer Analysis of Whole Genomes (PCAWG) project to develop a machine learning classifier able to accurately distinguish among 23 major cancer types using information derived from somatic mutations alone. This demonstrates the feasibility of automated cancer type discrimination based on next-generation sequencing of clinical samples. In addition, this work opens the possibility of determining the origin of tumours detected by the emerging technology of deep sequencing of circulating cell-free DNA in blood plasma.

cancer biology

Whole Genomes Define Concordance of Matched Primary, Xenograft, and Organoid Models of Pancreas Cancer

Pancreatic ductal adenocarcinoma (PDAC) has the worst prognosis among solid malignancies and improved therapeutic strategies are needed to improve outcomes. Patient-derived xenografts (PDX) and patient-derived organoids (PDO) serve as promising tools to identify new drugs with therapeutic potential in PDAC. For these preclinical disease models to be effective, they should both recapitulate the molecular heterogeneity of PDAC and validate patient-specific therapeutic sensitivities. To date however, deep characterization of PDAC PDX and PDO models and comparison with matched human tumour remains largely unaddressed at the whole genome level. We conducted a comprehensive assessment of the genetic landscape of 16 whole-genome pairs of tumours and matched PDX, from primary PDAC and liver metastasis, including a unique cohort of 5 trios of matched primary tumour, PDX, and PDO. We developed a new pipeline to score concordance between PDAC models and their paired human tumours for genomic events, including mutations, structural variations, and copy number variations. Comparison of genomic events in the tumours and matched disease models displayed single-gene concordance across major PDAC driver genes, and genome-wide similarities of copy number changes. Genome-wide and chromosome-centric analysis of structural variation (SV) events revealed high variability across tumours and disease models, but also highlighted previously unrecognized concordance across chromosomes that demonstrate clustered SV events. Our approach and results demonstrate that PDX and PDO recapitulate PDAC tumourigenesis with respect to simple somatic mutations and copy number changes, and capture major SV events that are found in both resected and metastatic tumours.

bioinformatics

Pan-cancer analysis of whole genomes

We report the integrative analysis of more than 2,600 whole cancer genomes and their matching normal tissues across 39 distinct tumour types. By studying whole genomes we have been able to catalogue non-coding cancer driver events, study patterns of structural variation, infer tumour evolution, probe the interactions among variants in the germline genome, the tumour genome and the transcriptome, and derive an understanding of how coding and non-coding variations together contribute to driving individual patient's tumours. This work represents the most comprehensive look at cancer whole genomes to date. NOTE TO READERS: This is an incomplete draft of the marker paper for the Pan-Cancer Analysis of Whole Genomes Project, and is intended to provide the background information for a series of in-depth papers that will be posted to BioRixv during the summer of 2017.

cancer biology

Large-Scale Uniform Analysis of Cancer Whole Genomes in Multiple Computing Environments

The International Cancer Genome Consortium (ICGC)s Pan-Cancer Analysis of Whole Genomes (PCAWG) project aimed to categorize somatic and germline variations in both coding and non-coding regions in over 2,800 cancer patients. To provide this dataset to the research working groups for downstream analysis, the PCAWG Technical Working Group marshalled ~800TB of sequencing data from distributed geographical locations; developed portable software for uniform alignment, variant calling, artifact filtering and variant merging; performed the analysis in a geographically and technologically disparate collection of compute environments; and disseminated high-quality validated consensus variants to the working groups. The PCAWG dataset has been mirrored to multiple repositories and can be located using the ICGC Data Portal. The PCAWG workflows are also available as Docker images through Dockstore enabling researchers to replicate our analysis on their own data.

genomics