Search bioRxivSearch

Biology subjects

Brandes, N.

Publications and source records attributed to Brandes, N..

2 recordsLinked to original sources

Substantial Batch Effects in TCGA Exome Sequences Undermine Pan-Cancer Analysis of Germline Variants

BackgroundIn recent years, research on cancer predisposition germline variants has emerged as a prominent field. The identity of somatic mutations is based on a reliable mapping of the patient germline variants. In addition, the statistics of germline variants frequencies in healthy individuals and cancer patients is the basis for seeking candidates for cancer predisposition genes. The Cancer Genome Atlas (TCGA) is one of the main sources of such data, providing a diverse collection of molecular data including deep sequencing for more than 30 types of cancer from >10,000 patients.\n\nMethodsOur hypothesis in this study is that whole exome sequences from healthy blood samples of cancer patients are not expected to show systematic differences among cancer types. To test this hypothesis, we analyzed common and rare germline variants across six cancer types, covering 2,241 samples from TCGA. In our analysis we accounted for inherent variables in the data including the different variant calling protocols, sequencing platforms, and ethnicity.\n\nResultsWe report on substantial batch effects in germline variants associated with cancer types. We attribute the effect to the specific sequencing centers that produced the data. Specifically, we measured 30% variability in the number of reported germline variants per sample across sequencing centers. The batch effect is further expressed in nucleotide composition and variant frequencies. Importantly, the batch effect causes substantial differences in germline variant distribution patterns across numerous genes, including prominent cancer predisposition genes such as BRCA1, RET, MAX, and KRAS. For most of known cancer predisposition genes, we found a distinct batch-dependent difference in germline variants.\n\nConclusionTCGA germline data is exposed to strong batch effects with substantial variabilities among TCGA sequencing centers. We claim that those batch effects are consequential for numerous TCGA pan-cancer studies. In particular, these effects may compromise the reliability and the potency to detect new cancer predisposition genes. Furthermore, interpretation of pan-cancer analyses should be revisited in view of the source of the genomic data after accounting for the reported batch effects.

cancer biology

Modeling Functional Genetic Alteration in Cancer Reveals New Candidate Driver Genes

Compiling the catalogue of genes actively involved in tumorigenesis (known as cancer drivers) is an ongoing endeavor, with profound implications to the understanding of tumorigenesis and treatment of the disease. An abundance of computational methods have been developed to screening the genome for candidate driver genes based on genomic data of somatic mutations in tumors. Most methods rely on detecting genes displaying excessive mutation rates compared to some background model. This approach is susceptible to false discoveries, due to its sensitivity to the assumptions of the background model, such as the need to account for hyper-mutated samples, cancer types and genomic loci. We present a fundamentally different approach. Instead of focusing on the number of mutations, we examine their content, and their expected effects on the functions of genes. We use a machine-learning model to predict functional effect scores of somatic mutations. For each gene, we compare the distribution of observed effect scores with the distribution expected at random, and report genes showing significant bias. By applying our framework on the ~20k protein-coding human genes, we detected 593 genes showing significant bias towards harmful mutations in the context of cancer. In contrast, we found only 6 significant genes biased in the opposite direction. The list of 593 genes, constructed without any prior knowledge of their role in cancer, shows an overwhelming overlap with known cancer driver genes, but also highlights many overlooked genes. These overlooked genes are promising candidates for novel cancer drivers. Our model is generic and is not restricted to the context of cancer. Applying the same framework to data of human-population genetic variation reveals the opposite trend. Unlike cancer, which is dominated by a bias towards harmful mutations, long-term evolution in healthy individuals results a bias towards less harmful mutations. The underlying assumptions of our framework are minimal, making it ideal for analyzing genetic data in search of genes subjected to positive or negative selection. It is fully open sourced and available for installation and use. Our framework presents a substantial development towards the application of state-of-the-art machine-learning algorithms in genetic studies.

bioinformatics