Search bioRxivSearch

Biology subjects

Karchin, R.

Publications and source records attributed to Karchin, R..

5 recordsLinked to original sources

Enhanced context reveals the scope of somatic missense mutations driving human cancers

Large-scale cancer sequencing studies of patient cohorts have statistically implicated many genes driving cancer growth and progression, and their identification has yielded substantial translational impact. However, a remaining challenge is to increase the resolution of driver prediction from the gene level to the mutation level, because mutation-level predictions are more closely aligned with the goal of precision cancer medicine. Here we present CHASMplus, a computational method, that is uniquely capable of identifying driver missense mutations, including those specific to a cancer type, as evidenced by significantly superior performance on diverse benchmarks. Applied to 8,657 tumor samples across 32 cancer types in The Cancer Genome Atlas, CHASMplus identifies over 4,000 unique driver missense mutations in 240 genes, supporting a prominent role for rare driver mutations. We show which TCGA cancer types are likely to yield discovery of new driver missense mutations by additional sequencing, which has important implications for public policy. SignificanceMissense mutations are the most frequent mutation type in cancers and the most difficult to interpret. While many computational methods have been developed to predict whether genes are cancer drivers or whether missense mutations are generally deleterious or pathogenic, there has not previously been a method to score the oncogenic impact of a missense mutation specifically by cancer type, limiting adoption of computational missense mutation predictors in the clinic. Cancer patients are routinely sequenced with targeted panels of cancer driver genes, but such genes contain a mixture of driver and passenger missense mutations which differ by cancer type. A patients therapeutic response to drugs and optimal assignment to a clinical trial depends on both the specific mutation in the gene of interest and cancer type. We present a new machine learning method honed for each TCGA cancer type, and a resource for fast lookup of the cancer-specific driver propensity of every possible missense mutation in the human exome.

bioinformatics

Non-invasive detection of upper tract urothelial carcinomas through the analysis of driver gene mutations and aneuploidy in urine

Upper tract urothelial carcinomas (UTUC) of the renal pelvis or ureter can be difficult to detect and challenging to diagnose. Here, we report the development and application of a non-invasive test for UTUC based on molecular analyses of DNA recovered from cells shed into the urine. The test, called UroSEEK, incorporates assays for mutations in eleven genes frequently mutated in urologic malignancies and for allelic imbalances on 39 chromosome arms. At least one genetic abnormality was detected in 75% of urinary cell samples from 56 UTUC patients but in only 0.5% of 188 samples from healthy individuals. The assay was considerably more sensitive than urine cytology, the current standard-of-care. UroSEEK therefore has the potential to be used for screening or to aid in diagnosis in patients at increased risk for UTUC, such as those exposed to herbal remedies containing the carcinogen aristolochic acid.

cancer biology

Non-invasive detection of bladder cancer through the analysis of driver gene mutations and aneuploidy

Current non-invasive approaches for bladder cancer (BC) detection are suboptimal. We report the development of non-invasive molecular test for BC using DNA recovered from cells shed into urine. This \"UroSEEK\" test incorporates assays for mutations in 11 genes and copy number changes on 39 chromosome arms. We first evaluated 570 urine samples from patients at risk for BC (microscopic hematuria or dysuria). UroSEEK was positive in 83% of patients that developed BC, but in only 7% of patients who did not develop BC. Combined with cytology, 95% of patients that developed BC were positive. We then evaluated 322 urine samples from patients soon after their BCs had been surgically resected. UroSEEK detected abnormalities in 66% of the urine samples from these patients, sometimes up to 4 years prior to clinical evidence of residual neoplasia, while cytology was positive in only 25% of such urine samples. The advantages of UroSEEK over cytology were particularly evident in low-grade tumors, wherein cytology detected none while UroSEEK detected 67% of 49 cases. These results establish the foundation for a new, non-invasive approach to the detection of BC in patients at risk for initial or recurrent disease.

cancer biology

CRAVAT 4: Cancer-Related Analysis of Variants Toolkit

Cancer sequencing studies are increasingly comprehensive and well-powered, returning long lists of somatic mutations that can be difficult to sort and interpret. Diligent analysis and quality control can require multiple computational tools of distinct utility and producing disparate output, creating additional challenges for the investigator. The Cancer-Related Analysis of Variants Toolkit (CRAVAT) is an evolving suite of informatics tools for mutation interpretation that includes mutation projecting and quality control, impact prediction and extensive annotation, gene- and mutation-level interpretation including joint prioritization of all nonsilent consequence types, and structural and mechanistic visualization. Results from CRAVAT submissions are explored in an interactive, user-friendly web-environment with dynamic filtering and sorting designed to highlight the most informative mutation, even in the context of very large studies. CRAVAT can be run on a public web-portal, in the cloud, or downloaded for local use, and is easily integrated with other methods for cancer omics analysis.\n\nConflict of interestAll authors declare no potential conflict of interest.

bioinformatics

Prediction of peptide binding to MHC Class I proteins in the age of deep learning

Binding of peptides to Major Histocompatibility Complex (MHC) proteins is a critical step in immune response. Peptides bound to MHCs are recognized by CD8+ (MHC Class I) and CD4+ (MHC Class II) T-cells. Successful prediction of which peptides will bind to specific MHC alleles would benefit many cancer immunotherapy appications. Currently, supervised machine learning is the leading computational approach to predict peptide-MHC binding, and a number of methods, trained using results of binding assays, have been published. Many clinical researchers are dissatisfied with the sensitivity and specificity of currently available methods and the limited number of alleles for which they can be applied. We evaluated several recent methods to predict peptide-MHC Class I binding affinities and a new method of our own design (MHCnuggets). We used a high-quality benchmark set of 51 alleles, which has been applied previously. The neural network methods NetMHC, NetMHCpan, MHCflurry, and MHCnuggets achieved similar best-in-class prediction performance in our testing, and of these methods MHCnuggets was significantly faster. MHCnuggets is a gated recurrent neural network, and the only method to our knowledge which can handle peptides of any length, without artificial lengthening and shortening. Seventeen alleles were problematic for all tested methods. Prediction difficulties could be explained by deficiencies in the training and testing examples in the benchmark, suggesting that biological differences in allele-specific binding properties are not as important as previously claimed. Advances in accuracy and speed of computational methods to predict peptide-MHC affinity are urgently needed. These methods will be at the core of pipelines to identify patients who will benefit from immunotherapy, based on tumor-derived somatic mutations. Machine learning methods, such as MHCnuggets, which efficiently handle peptides of any length will be increasingly important for the challenges of predicting immunogenic response for MHC Class II alleles.\n\nAuthor SummaryMachine learning methods are a popular approach for predicting whether a peptide will bind to Major Histocompatibility Complex (MHC) proteins, a critical step in activation of cytotoxic T-cells. The input to these methods is a peptide sequence and an MHC allele of interest, and the output is the predicted binding affinity. MHC Class I and II proteins bind peptides of 8-11 amino acids and 16-26 amino acids respectively. This has been an obstacle for machine learning, because the methods used to date can only handle fixed-length inputs. We show that a recently developed technique known as gated recurrent neural networks can handle peptides of variable length and predict peptide-MHC binding as well or better than existing methods, at substantially faster speeds. Our results have implications for the hundreds of MHC alleles that cannot be predicted with current methods.

bioinformatics