Search bioRxivSearch

Biology subjects

Boutros, P. C.

Publications and source records attributed to Boutros, P. C..

17 recordsLinked to original sources

The Inter and Intra-Tumoural Heterogeneity of Subclonal Reconstruction

Whole-genome sequencing can be used to estimate subclonal populations in tumours and this intra-tumoural heterogeneity is linked to clinical outcomes. Many algorithms have been developed for subclonal reconstruction, but their variabilities and consistencies are largely unknown. We evaluated sixteen pipelines for reconstructing the evolutionary histories of 293 localized prostate cancers from single samples, and eighteen pipelines for the reconstruction of 10 tumours with multi-region sampling. We show that predictions of subclonal architecture and timing of somatic mutations vary extensively across pipelines. Pipelines show consistent types of biases, with those incorporating SomaticSniper and Battenberg preferentially predicting homogenous cancer cell populations and those using MuTect tending to predict multiple populations of cancer cells. Subclonal reconstructions using multi-region sampling confirm that single-sample reconstructions systematically underestimate intra-tumoural heterogeneity, predicting on average fewer than half of the cancer cell populations identified by multi-region sequencing. Overall, these biases suggest caution in interpreting specific architectures and subclonal variants.

cancer biology

HPCI: A Perl module for writing cluster-portable bioinformatics pipelines

BackgroundMost biocomputing pipelines are run on clusters of computers. Each type of cluster has its own API (application programming interface). That API defines how a program that is to run on the cluster must request the submission, content and monitoring of jobs to be run on the cluster. Sometimes, it is desirable to run the same pipeline on different types of cluster. This can happen in situations including when:\n\nO_LIdifferent labs are collaborating, but they do not use the same type of cluster\nC_LIO_LIa pipeline is released to other labs as open source or commercial software\nC_LIO_LIa lab has access to multiple types of cluster, and wants to choose between them for scaling, cost or other purposes\nC_LIO_LIa lab is migrating their infrastructure from one cluster type to another\nC_LIO_LIduring testing or travelling, it is often desired to run on a single computer\nC_LI\n\nHowever, since each type of cluster has its own API, code that runs jobs on one type of cluster needs to be re-written if it is desired to run that application on a different type of cluster. To resolve this problem, we created a software module to generalize the submission of pipelines across computing environments, including local compute, clouds and clusters.\n\nResultsHPCI (High Performance Computing Interface) is a Perl module that provides the interface to a standardized generic cluster.\n\nWhen the HPCI module is used, it accepts a parameter to specify the cluster type. The HPCI module uses this to load a driver HPCD:: . This is used to translate the abstract HPCI interface to the specific software interface.\n\nSimply by changing the cluster parameter, the same pipeline can be run on a different type of cluster with no other changes.\n\nConclusionThe HPCI module assists in writing Perl programs that can be run in different lab environments, with different site configuration requirements and different types of hardware clusters. Rather than having to re-write portions of the program, it is only necessary to change a configuration file.\n\nUsing HPCI, an application can manage collections of jobs to be runs, specify ordering dependencies, detect success or failure of jobs run and allow automatic retry of failed jobs (allowing for the possibility of a changed configuration such as when the original attempt specified an inadequate memory allotment).

bioinformatics

Integrative pathway enrichment analysis of multivariate omics data

Multi-omics datasets quantify complementary aspects of molecular biology and thus pose challenges to data interpretation and hypothesis generation. ActivePathways is an integrative method that discovers significantly enriched pathways across multiple omics datasets using a statistical data fusion approach, rationalizes contributing evidence and highlights associated genes. We demonstrate its utility by analyzing coding and non-coding mutations from 2,583 whole cancer genomes, revealing frequently mutated hallmark pathways and a long tail of known and putative cancer driver genes. We also studied prognostic molecular pathways in breast cancer subtypes by integrating genomic and transcriptomic features of tumors and tumor-adjacent cells and found significant associations with immune response processes and anti-apoptotic signaling pathways. ActivePathways is a versatile method that improves systems-level understanding of cellular organization in health and disease through integration of multiple molecular datasets and pathway annotations.

bioinformatics

Accurate Reference-Free Somatic Variant-Calling by Integrating Genomic, Sequencing and Population Data

The detection of somatic single nucleotide variants (SNVs) is critical in both research and clinical applications. Studies of human cancer typically use matched normal (reference) samples from a distant tissue to increase SNV prediction accuracy. This process both doubles sequencing costs and poses challenges when reference samples are not readily available, such as for many cell-lines. To address these challenges, we created S22S: an approach for the prediction of somatic mutations without need for matched reference tissue. S22S takes underlying sequence data, augments them with genomic background context and population frequency information, and classifies SNVs as somatic or non-somatic. We validated S22S using primary tumor/normal pairs from four tumor types, spanning two different sequencing technologies. S22S robustly identifies somatic SNVs, with the area under the precision recall curve reaching 0.97 in kidney clear cell carcinoma, comparable to the best tumor/normal analysis pipelines. S22S is freely available at http://labs.oicr.on.ca/Boutros-lab/software/s22s.

bioinformatics

Modeling the MYC-driven normal-to-tumour switch in breast cancer.

The potent MYC oncoprotein is deregulated in many human cancers, including breast carcinoma, and is associated with aggressive disease. To understand the mechanisms and vulnerabilities of MYC-driven breast cancer, we have generated an in vivo model that mimics human disease in response to MYC deregulation. MCF10A cells ectopically expressing a common breast cancer mutation in the PI3 kinase pathway (PIK3CAH1047R) lead to the development of organized acinar structures in mice. However, expressing both PIK3CAH1047R and deregulated-MYC lead to the development of invasive ductal carcinoma, thus creating a model in which a MYC-dependent normal-to-tumour switch occurs in vivo. These MYC-driven tumors exhibit classic hallmarks of human breast cancer at both the pathological and molecular levels. Moreover, tumour growth is dependent upon sustained deregulated MYC expression, further demonstrating addiction to this potent oncogene and regulator of gene transcription. We therefore provide a MYC-dependent human model of breast cancer which can be assayed for in vivo tumour initiation, proliferation, and transformation from normal breast acini into invasive breast carcinoma. Taken together, we anticipate that this novel MYC-driven transformation model will be a useful research tool to both better understand MYCs oncogenic function and identify therapeutic vulnerabilities.

cancer biology

Fast Nonnegative Matrix Factorization andApplications to Pattern Extraction, Deconvolutionand Imputation

Nonnegative matrix factorization (NMF) is a technique widely used in various fields, including artificial intelligence (AI), signal processing and bioinformatics. However existing algorithms and R packages cannot be applied to large matrices due to their slow convergence, and cannot handle missing values. In addition, most NMF research focuses only on blind decompositions: decomposition without utilizing prior knowledge. We adapt the idea of sequential coordinate-wise descent to NMF to increase the convergence rate. Our NMF algorithm thus handles missing values naturally and integrates prior knowledge to guide NMF towards a more meaningful decomposition. To support its use, we describe a novel imputation-based method to determine the rank of decomposition. All our algorithms are implemented in the R package NNLM, which is freely available on CRAN.

bioinformatics

Portraits of genetic intra-tumour heterogeneity and subclonal selection across cancer types

Intra-tumor heterogeneity (ITH) is a mechanism of therapeutic resistance and therefore an important clinical challenge. However, the extent, origin and drivers of ITH across cancer types are poorly understood. To address this question, we extensively characterize ITH across whole-genome sequences of 2,658 cancer samples, spanning 38 cancer types. Nearly all informative samples (95.1%) contain evidence of distinct subclonal expansions, with frequent branching relationships between subclones. We observe positive selection of subclonal driver mutations across most cancer types, and identify cancer type specific subclonal patterns of driver gene mutations, fusions, structural variants and copy-number alterations, as well as dynamic changes in mutational processes between subclonal expansions. Our results underline the importance of ITH and its drivers in tumor evolution, and provide an unprecedented pan-cancer resource of comprehensively annotated subclonal events from whole-genome sequencing data.

cancer biology

Creating Standards for Evaluating Tumour Subclonal Reconstruction

Tumours evolve through time and space. Computational techniques have been developed to infer their evolutionary dynamics from DNA sequencing data. A growing number of studies have used these approaches to link molecular cancer evolution to clinical progression and response to therapy. There has not yet been a systematic evaluation of methods for reconstructing tumour subclonality, in part due to the underlying mathematical and biological complexity and to difficulties in creating gold-standards. To fill this gap, we systematically elucidated the key algorithmic problems in subclonal reconstruction and developed mathematically valid quantitative metrics for evaluating them. We then created approaches to simulate realistic tumour genomes, harbouring all known mutation types and processes both clonally and subclonally. We then simulated 580 tumour genomes for reconstruction, varying tumour read-depth and benchmarking somatic variant detection and subclonal reconstruction strategies. The inference of tumour phylogenies is rapidly becoming standard practice in cancer genome analysis; this study creates a baseline for its evaluation.

bioinformatics

Subnetwork-based prognostic biomarkers exhibit performance and robustness superior to gene-based biomarkers in breast cancer

BackgroundEffective classification of cancer patients into groups with differential survival remains an important and unsolved challenge. Biomarkers have been developed based on mRNA abundance data, but their replicability and clinical utility is modest. Integrating functional information, such as pathway data, has been suggested to improve biomarker performance. To date, however, the advantages of subnetwork-based biomarkers have not been quantified.\n\nResultsWe deeply sampled the population of prognostic gene-based and subnetwork-based biomarkers in a breast cancer meta-dataset of 4,960 patients. Analysing the performance and robustness of 22,000,000 gene biomarkers and 6,250,000 subnetwork biomarkers across twenty different training:testing cohort partitions of the meta-dataset revealed that subnetwork biomarkers exhibit superior overall performance and higher concordance across partitions. We find evidence of an upper bound for optimal biomarker size of [~]200 genes or [~]100 subnetworks. Additionally, with both biomarker feature types, larger biomarkers tend to show less consistency in performance across partitions, suggestive of over-fitting. Finally, an evaluation of varying training cohort sizes quantifies the effects of training cohort size.\n\nConclusionsMany groups are developing techniques for exploiting network-based representations of biological pathways to characterize cancer and other diseases. By considering the distribution of gene- and subnetwork-based biomarkers, we show that pathway data improves performance and replicability, and that smaller biomarkers are more robust across patient cohorts. These insights may facilitate development of clinically useful biomarkers.

bioinformatics

Network-Based Biomarkers Enable Cross-Disease Biomarker Discovery

Biomarkers lie at the heart of precision medicine, biodiversity monitoring, agricultural pathogen detection, amongst others. Surprisingly, while rapid genomic profiling is becoming ubiquitous, the development of biomarkers almost always involves the application of bespoke techniques that cannot be directly applied to other datasets. There is an urgent need for a systematic methodology to create biologically-interpretable molecular models that robustly predict key phenotypes. We therefore created SIMMS: an algorithm that fragments pathways into functional modules and uses these to predict phenotypes. We applied SIMMS to multiple data-types across four diseases, and in each it reproducibly identified subtypes, made superior predictions to the best bespoke approaches, and identified known and novel signaling nodes. To demonstrate its ability on a new dataset, we measured 33 genes/nodes of the PIK3CA pathway in 1,734 FFPE breast tumours and created a four-subnetwork prediction model. This model significantly out-performed existing clinically-used molecular tests in an independent 1,742-patient validation cohort. SIMMS is generic and can work with any molecular data or biological network, and is freely available at: https://cran.r-project.org/web/packages/SIMMS.

bioinformatics

The Origins and Consequences of Localized and Global Somatic Hypermutation

Cancer is a disease of the genome, but the dramatic inter-patient variability in mutation number is poorly understood. Tumours of the same type can differ by orders of magnitude in their mutation rate. To understand potential drivers and consequences of the underlying heterogeneity in mutation rate across tumours, we evaluated both local and global measures of mutation density: both single-stranded and double-stranded DNA breaks in 2,460 tumours of 38 cancer types. We find that SCNAs in thousands of genes are associated with elevated rates of point-mutations, while similarly point-mutation patterns in dozens of genes are associated with specific patterns of DNA double-stranded breaks. These candidate drivers of mutation density are enriched for known cancer drivers, and preferentially occur early in tumour evolution, appearing clonally in all cells of a tumour. To supplement this understanding of global mutation density, we developed and validated a tool called SeqKat to identify localized \"rainstorms\" of point-mutations (kataegis). We show that rates of kataegis differ by four orders of magnitude across tumour types, with malignant lymphomas showing the highest. Tumours with TP53 mutations were 2.6-times more likely to harbour a kataegic event than those without, and 239 SCNAs were associated with elevated rates of kataegis, including loss of the tumour-suppressor CDKN2A. We identify novel subtypes of kataegic events not associated with aberrant APOBEC activity, and find that these are localized to specific cellular regions, enriched for MYC-target genes. Kataegic events were associated with patient survival in some, but not all tumour types, highlighting a combination of global and tumour-type specific effects. Taken together, we reveal a landscape of genes driving localized and tumour-specific hyper-mutation, and reveal novel mutational processes at play in specific tumour types.

cancer biology

Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes

Assessing accuracy of NGS variant calling is immensely facilitated by a robust benchmarking strategy and tools to carry it out in a standard way. Benchmarking variant calls requires careful attention to definitions of performance metrics, sophisticated comparison approaches, and stratification by variant type and genome context. The Global Alliance for Genomics and Health (GA4GH) Benchmarking Team has developed standardized performance metrics and tools for benchmarking germline small variant calls. This team includes representatives from sequencing technology developers, government agencies, academic bioinformatics researchers, clinical laboratories, and commercial technology and bioinformatics developers for whom benchmarking variant calls is essential to their work. Benchmarking variant calls is a challenging problem for many reasons:\n\nO_LIEvaluating variant calls requires complex matching algorithms and standardized counting because the same variant may be represented differently in truth and query callsets.\nC_LIO_LIDefining and interpreting resulting metrics such as precision (aka positive predictive value = TP/(TP+FP)) and recall (aka sensitivity = TP/(TP+FN)) requires standardization to draw robust conclusions about comparative performance for different variant calling methods.\nC_LIO_LIPerformance of NGS methods can vary depending on variant types and genome context; and as a result understanding performance requires meaningful stratification.\nC_LIO_LIHigh-confidence variant calls and regions that can be used as \"truth\" to accurately identify false positives and negatives are difficult to define, and reliable calls for the most challenging regions and variants remain out of reach.\nC_LI\n\nWe have made significant progress on standardizing comparison methods, metric definitions and reporting, as well as developing and using truth sets. Our methods are publicly available on GitHub (https://github.com/ga4gh/benchmarking-tools) and in a web-based app on precisionFDA, which allow users to compare their variant calls against truth sets and to obtain a standardized report on their variant calling performance. Our methods have been piloted in the precisionFDA variant calling challenges to identify the best-in-class variant calling methods within high-confidence regions. Finally, we recommend a set of best practices for using our tools and critically evaluating the results.

genomics

Valection: Design Optimization for Validation and Verification Studies

BackgroundPlatform-specific error profiles necessitate confirmatory studies where predictions made on data generated using one technology are additionally verified by processing the same samples on an orthogonal technology. In disciplines that rely heavily on high-throughput data generation, such as genomics, reducing the impact of false positive and false negative rates in results is a top priority. However, verifying all predictions can be costly and redundant, and testing a subset of findings is often used to estimate the true error profile. To determine how to create subsets of predictions for validation that maximize inference of global error profiles, we developed Valection, a software program that implements multiple strategies for the selection of verification candidates.\n\nResultsTo evaluate these selection strategies, we obtained 261 sets of somatic mutation calls from a single-nucleotide variant caller benchmarking challenge where 21 teams competed on whole-genome sequencing datasets of three computationally-simulated tumours. By using synthetic data, we had complete ground truth of the tumours mutations and, therefore, we were able to accurately determine how estimates from the selected subset of verification candidates compared to the complete prediction set. We found that selection strategy performance depends on several verification study characteristics. In particular the verification budget of the experiment (i.e. how many candidates can be selected) is shown to influence estimates.\n\nConclusionsThe Valection framework is flexible, allowing for the implementation of additional selection algorithms in the future. Its applicability extends to any discipline that relies on experimental verification and will benefit from the optimization of verification candidate selection.

bioinformatics

Candidate cancer driver mutations in super-enhancers and long-range chromatin interaction networks

A comprehensive catalogue of the mutations that drive tumorigenesis and progression is essential to understanding tumor biology and developing therapies. Protein-coding driver mutations have been well-characterized by large exome-sequencing studies, however many tumors have no mutations in protein-coding driver genes. Non-coding mutations are thought to explain many of these cases, however few non-coding drivers besides TERT promoter are known. To fill this gap, we analyzed 150,000 cis-regulatory regions in 1,844 whole cancer genomes from the ICGC-TCGA PCAWG project. Using our new method, ActiveDriverWGS, we found 41 frequently mutated regulatory elements (FMREs) enriched in non-coding SNVs and indels (FDR<0.05) characterized by aging-associated mutation signatures and frequent structural variants. Most FMREs are distal from genes, reported here for the first time and also recovered by additional driver discovery methods. FMREs were enriched in super-enhancers, H3K27ac enhancer marks of primary tumors and long-range chromatin interactions, suggesting that the mutations drive cancer by distally controlling gene expression through threedimensional genome organization. In support of this hypothesis, the chromatin interaction network of FMREs and target genes revealed associations of mutations and differential gene expression of known and novel cancer genes (e.g., CNNB1IP1, RCC1), activation of immune response pathways and altered enhancer marks. Thus distal genomic regions may include additional, infrequently mutated drivers that act on target genes via chromatin loops. Our study is an important step towards finding such regulatory regions and deciphering the somatic mutation landscape of the non-coding genome.

cancer biology

Combining accurate tumour genome simulation with crowd sourcing to benchmark somatic structural variant detection

BackgroundThe phenotypes of cancer cells are driven in part by somatic structural variants. Structural variants can initiate tumors, enhance their aggressiveness and provide unique therapeutic opportunities. Whole-genome sequencing of tumors can allow exhaustive identification of the specific structural variants present in an individual cancer, facilitating both clinical diagnostics and the discovery of novel mutagenic mechanisms. A plethora of somatic structural variant detection algorithms have been created to enable these discoveries, however there are no systematic benchmarks of them. Rigorous performance evaluation of somatic structural variant detection methods has been challenged by the lack of gold-standards, extensive resource requirements and difficulties arising from the need to share personal genomic information.\n\nResultsTo facilitate structural variant detection algorithm evaluations, we create a robust simulation framework for somatic structural variants by extending the BAMSurgeon algorithm. We then organize and enable a crowd-sourced benchmarking within the ICGC-TCGA DREAM Somatic Mutation Calling Challenge (SMC-DNA). We report here the results of structural variant benchmarking on three different tumors, comprising 204 submissions from 15 teams. In addition to ranking methods, we identify characteristic error-profiles of individual algorithms and general trends across them. Surprisingly, we find that ensembles of analysis pipelines do not always outperform the best individual method, indicating a need for new ways to aggregate somatic structural variant detection approaches.\n\nConclusionsThe synthetic tumors and somatic structural variant detection leaderboards remain available as a community benchmarking resource, and BAMSurgeon is available at https://github.com/adamewing/bamsurgeon.

bioinformatics

Germline Contamination and Leakage in Whole Genome Somatic Single Nucleotide Variant Detection

BackgroundThe clinical sequencing of cancer genomes to personalize therapy is becoming routine across the world. However, concerns over patient re-identification from these data lead to questions about how tightly access should be controlled. It is not thought to be possible to re-identify patients from somatic variant data. However, somatic variant detection pipelines can mistakenly identify germline variants as somatic ones, a process called \"germline leakage\". The rate of germline leakage across different somatic variant detection pipelines is not well-understood, and it is uncertain whether or not somatic variant calls should be considered re-identifiable. To fill this gap, we quantified germline leakage across 259 sets of whole-genome somatic single nucleotide variant (SNVs) predictions made by 21 teams as part of the ICGC-TCGA DREAM Somatic Mutation Calling Challenge.\n\nResultsThe median somatic SNV prediction set contained 4,325 somatic SNVs and leaked one germline polymorphism. The level of germline leakage was inversely correlated with somatic SNV prediction accuracy and positively correlated with the amount of infiltrating normal cells. The specific germline variants leaked differed by tumour and algorithm. To aid in quantitation and correction of leakage, we created a tool, called GermlineFilter, for use in public-facing somatic SNV databases.\n\nConclusionsThe potential for patient re-identification from leaked germline variants in somatic SNV predictions has led to divergent open data access policies, based on different assessments of the risks. Indeed, a single, well-publicized re-identification event could reshape public perceptions of the values of genomic data sharing. We find that modern somatic SNV prediction pipelines have low germline-leakage rates, which can be further reduced, especially for cloud-sharing, using pre-filtering software.

genomics

The evolutionary history of 2,658 cancers

Cancer develops through a process of somatic evolution. Here, we use whole-genome sequencing of 2,778 tumour samples from 2,658 donors to reconstruct the life history, evolution of mutational processes, and driver mutation sequences of 39 cancer types. The early phases of oncogenesis are driven by point mutations in a small set of driver genes, often including biallelic inactivation of tumour suppressors. Early oncogenesis is also characterised by specific copy number gains, such as trisomy 7 in glioblastoma or isochromosome 17q in medulloblastoma. By contrast, increased genomic instability, a nearly four-fold diversification of driver genes, and an acceleration of point mutation processes are features of later stages. Copy-number alterations often occur in mitotic crises leading to simultaneous gains of multiple chromosomal segments. Timing analysis suggests that driver mutations often precede diagnosis by many years, and in some cases decades, providing a window of opportunity for early cancer detection.

cancer biology