Search bioRxivSearch

Biology subjects

Fang, L. T.

Publications and source records attributed to Fang, L. T..

2 recordsLinked to original sources

Robust Cancer Mutation Detection with Deep Learning Models Derived from Tumor-Normal Sequencing Data

Accurate detection of somatic mutations is challenging but critical to the understanding of cancer formation, progression, and treatment. We recently proposed NeuSomatic, the first deep convolutional neural network based somatic mutation detection approach and demonstrated performance advantages on in silico data. In this study, we used the first comprehensive and well-characterized somatic reference samples from the SEQC-II consortium to investigate best practices for utilizing deep learning framework in cancer mutation detection. Using the high-confidence somatic mutations established for these reference samples by the consortium, we identified strategies for building robust models on multiple datasets derived from samples representing real scenarios. The proposed strategies achieved high robustness across multiple sequencing technologies such as WGS, WES, AmpliSeq target sequencing for fresh and FFPE DNA input, varying tumor/normal purities, and different coverages (ranging from 10x - 2000x). NeuSomatic significantly outperformed conventional detection approaches in general, as well as in challenging situations such as low coverage, low mutation frequency, DNA damage, and difficult genomic regions.

genomics

Establishing reference samples for detection of somatic mutations and germline variants with NGS technologies

We characterized two reference samples for NGS technologies: a human triple-negative breast cancer cell line and a matched normal cell line. Leveraging several whole-genome sequencing (WGS) platforms, multiple sequencing replicates, and orthogonal mutation detection bioinformatics pipelines, we minimized the potential biases from sequencing technologies, assays, and informatics. Thus, our \"truth sets\" were defined using evidence from 21 repeats of WGS runs with coverages ranging from 50X to 100X (a total of 140 billion reads). These \"truth sets\" present many relevant variants/mutations including 193 COSMIC mutations and 9,016 germline variants from the ClinVar database, nonsense mutations in BRCA1/2 and missense mutations in TP53 and FGFR1. Independent validation in three orthogonal experiments demonstrated a successful stress test of the truth set. We expect these reference materials and \"truth sets\" to facilitate assay development, qualification, validation, and proficiency testing. In addition, our methods can be extended to establish new fully characterized reference samples for the community.

bioinformatics