Search bioRxivSearch

Biology subjects

Venner, E.

Publications and source records attributed to Venner, E..

2 recordsLinked to original sources

Atlas-CNV: a validated approach to call Single-Exon CNVs in the eMERGESeq gene panel

PurposeTo provide a validated method to confidently identify exon-containing copy number variants (CNVs), with a low false discovery rate (FDR), in targeted sequencing data from a clinical laboratory with particular focus on single-exon CNVs.\n\nMethodsDNA sequence coverage data are normalized within each sample and subsequently exonic CNVs are identified in a batch of samples (midpool), when the target log2 ratio of the sample to the batch median exceeds defined thresholds. The quality of exonic CNV calls is assessed by C-scores (Z-like scores) using thresholds derived from gold standard samples and simulation studies. We integrate an ExonQC threshold to lower FDR and compare performance with alternate software (VisCap).\n\nResultsThirteen CNVs were used as a truth set to validate Atlas-CNV and compared with VisCap. We demonstrated FDR reduction in validation, simulation and 10,926 eMERGESeq samples without sensitivity loss. Sixty-four multi-exon and 29 single-exon CNVs with high C-scores were assessed by MLPA.\n\nConclusionsAtlas-CNV is validated as a method to identify exonic CNVs in targeted sequencing data generated in the clinical laboratory. The ExonQC and C-score assignment can reduce FDR (identification of targets with high variance) and improve calling accuracy of single-exon CNVs respectively. We proposed guidelines and criteria to identify high confidence single-exon CNVs.

genomics

xAtlas: Scalable small variant calling across heterogeneous next-generation sequencing experiments

MotivationThe rapid development of next-generation sequencing (NGS) technologies has lowered the barriers to genomic data generation, resulting in millions of samples sequenced across diverse experimental designs. The growing volume and heterogeneity of these sequencing data complicate the further optimization of methods for identifying DNA variation, especially considering that curated highconfidence variant call sets commonly used to evaluate these methods are generally developed by reference to results from the analysis of comparatively small and homogeneous sample sets.\n\nResultsWe have developed xAtlas, an application for the identification of single nucleotide variants (SNV) and small insertions and deletions (indels) in NGS data. xAtlas is easily scalable and enables execution and retraining with rapid development cycles. Generation of variant calls in VCF or gVCF format from BAM or CRAM alignments is accomplished in less than one CPU-hour per 30x short-read human whole-genome. The retraining capabilities of xAtlas allow its core variant evaluation models to be optimized on new sample data and user-defined truth sets. Obtaining SNV and indels calls from xAtlas can be achieved more than 40 times faster than established methods while retaining the same accuracy.\n\nAvailabilityFreely available under a BSD 3-clause license at https://github.com/jfarek/xatlas.\n\nContactfarek@bcm.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics