bioRxiv · 10.1101/2020.06.05.136549
Mis-annotated multi nucleotide variants in public cancer genomics datasets can lead to inaccurate mutation calls with significant implications
Abstract
BackgroundNext generation sequencing is widely used in cancer to profile tumors and detect variants. Most somatic variant callers used in these pipelines identify variants at the lowest possible granularity - single nucleotide variants (SNVs). As a result, multiple adjacent SNVs are called individually instead of as a multi-nucleotide variant (MNV). The problem with this level of granularity is that the amino acid change from the individual SNVs within a codon could be different from the amino acid change based on the MNV that results from combining the SNVs. Most variant annotation tools do not account for this, leading to incorrect conclusions about the downstream effects of the variants. MethodHere, we used Variant Call Files (VCFs) from the TCGA Mutect2 caller, and developed a solution to merge SNVs to MNVs. Our custom script takes the phasing information from the SNV VCFs and based on a gene model, determines if SNVs are at the same codon and need to be merged into a MNV prior to variant annotation. ResultsWe analyzed 10,383 VCFs from TCGA and found 12,141 MNVs that were incorrectly annotated. Strikingly, the analysis of seven commonly mutated genes from 178 studies from cBioPortal revealed that MNVs were consistently missed in 20 of these studies, while they were correctly annotated in 15 more recent studies. The best and most common example of MNVs was found at the BRAF V600 locus, where several public datasets reported separate BRAF V600E and BRAF V600M variants, instead of a single merged V600K variant. ConclusionWhile some datasets merged MNVs correctly, many public datasets have not been corrected for this problem. As a best practice for variant calling, we recommend that MNVs be accounted for in NGS processing pipelines, thus improving analyses on the impact of somatic variants in cancer genomics.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Srinivasan, S., Kalinava, N., Aldana, R., Li, Z., Van Hagen, S., Rodenburg, S. Y. A., Wind-Rotolo, M., Sasson, A. S., Tang, H., Qian, X., Kirov, S.. 2020-06-06. Mis-annotated multi nucleotide variants in public cancer genomics datasets can lead to inaccurate mutation calls with significant implications. https://doi.org/10.1101/2020.06.05.136549
Cite the original work for its findings. Save a collection to share your selection of sources.