Search bioRxivSearch

Biology subjects

Salit, M.

Publications and source records attributed to Salit, M..

7 recordsLinked to original sources

genomeview - an extensible python-based genomics visualization engine

Visual inspection and analysis is integral to quality control, hypothesis generation, methods development and validation of genomic data. The richness and complexity of genomic data necessitates customized visualizations highlighting specific features of interest while hiding the often vast tide of irrelevant attributes. However, the majority of genome-visualization occurs either in general-purpose tools such as IGV (Robinson et al, 2011) or the UCSC Genome Browser (Kent et al, 2002) - which offer many options to adjust visualization parameters, but very little in the way of extensibility - or narrowly-focused tools aiming to solve a single visualization problem. Here, we present genomeview, a python-based visualization engine which is easy to extend and simple to integrate into existing analysis pipelines.

genomics

A Rigorous Interlaboratory Examination of the Need to Confirm NGS-Detected Variants by an Orthogonal Method in Clinical Genetic Testing

Orthogonal confirmation of NGS-detected germline variants has been standard practice, although published studies have suggested that confirmation of the highest quality calls may not always be necessary. The key question is how laboratories can establish criteria that consistently identify those NGS calls that require confirmation. Most prior studies addressing this question have limitations: These studies are generally small, omit statistical justification, and explore limited aspects of the underlying data. The rigorous definition of criteria that separate high-accuracy NGS calls from those that may or may not be true remains a critical issue.\n\nWe analyzed five reference samples and over 80,000 patient specimens from two laboratories. We examined quality metrics for approximately 200,000 NGS calls with orthogonal data, including 1662 false positives. A classification algorithm used these data to identify a battery of criteria that flag 100% of false positives as requiring confirmation (CI lower bound: 98.5-99.8% depending on variant type) while minimizing the number of flagged true positives. These criteria identify false positives that the previously published criteria miss. Sampling analysis showed that smaller datasets resulted in less effective criteria.\n\nOur methodology for determining test and laboratory-specific criteria can be generalized into a practical approach that can be used by many laboratories to help reduce the cost and time burden of confirmation without impacting clinical accuracy.

genetics

Reproducing bench-scale cell growth and productivity

Reproducing, exchanging, comparing, and building on each others work is foundational to technology advances.1 Advancing biotechnology calls for reliable reuse of engineered strains.2 Reliable reuse of engineered strains requires reproducible growth and productivity. To demonstrate reproducibility for biotechnology, we identified the experimental factors that have the greatest effect on the growth and productivity of our engineered strains.3-6 We present a draft of a Minimum Information Standard for Engineered Organism Experiments (MIEO) based on this method. We evaluated the effect of 22 factors on Escherichia coli (E. coli) engineered to produce the small molecule lycopene, and 18 factors on E. coli engineered to produce red fluorescent protein (RFP). Container geometry and shaking had the greatest effect on product titer and yield. We reproduced our results under two different conditions of reproducibility:7 conditions of use (different fractional factorial experiments), and time (48 biological replicates performed on 12 different days over four months).

synthetic biology

Reproducible integration of multiple sequencing datasets to form high-confidence SNP, indel, and reference calls for five human genome reference materials

Benchmark small variant calls from the Genome in a Bottle Consortium (GIAB) for the CEPH/HapMap genome NA12878 (HG001) have been used extensively for developing, optimizing, and demonstrating performance of sequencing and bioinformatics methods. Here, we develop a reproducible, cloud-based pipeline to integrate multiple sequencing datasets and form benchmark calls, enabling application to arbitrary human genomes. We use these reproducible methods to form high-confidence calls with respect to GRCh37 and GRCh38 for HG001 and 4 additional broadly-consented genomes from the Personal Genome Project that are available as NIST Reference Materials. These new genomes broad, open consent with few restrictions on availability of samples and data is enabling a uniquely diverse array of applications. Our new methods produce 17% more high-confidence SNPs, 176% more indels, and 12% larger regions than our previously published calls. To demonstrate that these calls can be used for accurate benchmarking, we compare other high-quality callsets to ours (e.g., Illumina Platinum Genomes), and we demonstrate that the majority of discordant calls are errors in the other callsets, We also highlight challenges in interpreting performance metrics when benchmarking against imperfect high-confidence calls. We show that benchmarking tools from the Global Alliance for Genomics and Health can be used with our calls to stratify performance metrics by variant type and genome context and elucidate strengths and weaknesses of a method.

genomics

Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes

Assessing accuracy of NGS variant calling is immensely facilitated by a robust benchmarking strategy and tools to carry it out in a standard way. Benchmarking variant calls requires careful attention to definitions of performance metrics, sophisticated comparison approaches, and stratification by variant type and genome context. The Global Alliance for Genomics and Health (GA4GH) Benchmarking Team has developed standardized performance metrics and tools for benchmarking germline small variant calls. This team includes representatives from sequencing technology developers, government agencies, academic bioinformatics researchers, clinical laboratories, and commercial technology and bioinformatics developers for whom benchmarking variant calls is essential to their work. Benchmarking variant calls is a challenging problem for many reasons:\n\nO_LIEvaluating variant calls requires complex matching algorithms and standardized counting because the same variant may be represented differently in truth and query callsets.\nC_LIO_LIDefining and interpreting resulting metrics such as precision (aka positive predictive value = TP/(TP+FP)) and recall (aka sensitivity = TP/(TP+FN)) requires standardization to draw robust conclusions about comparative performance for different variant calling methods.\nC_LIO_LIPerformance of NGS methods can vary depending on variant types and genome context; and as a result understanding performance requires meaningful stratification.\nC_LIO_LIHigh-confidence variant calls and regions that can be used as \"truth\" to accurately identify false positives and negatives are difficult to define, and reliable calls for the most challenging regions and variants remain out of reach.\nC_LI\n\nWe have made significant progress on standardizing comparison methods, metric definitions and reporting, as well as developing and using truth sets. Our methods are publicly available on GitHub (https://github.com/ga4gh/benchmarking-tools) and in a web-based app on precisionFDA, which allow users to compare their variant calls against truth sets and to obtain a standardized report on their variant calling performance. Our methods have been piloted in the precisionFDA variant calling challenges to identify the best-in-class variant calling methods within high-confidence regions. Finally, we recommend a set of best practices for using our tools and critically evaluating the results.

genomics

An interlaboratory study of complex variant detection

Next-generation sequencing (NGS) is widely used and cost-effective. Depending on the specific methods, NGS can have limitations detecting certain technically challenging variant types even though they are both prevalent in patients and medically important. These types are underrepresented in validation studies, hindering the uniform assessment of test methodologies by laboratory directors and clinicians. Specimens containing such variants can be difficult to obtain; thus, we evaluated a novel solution to this problem in which a diverse set of technically challenging variants was synthesized and introduced into a known genomic background. This specimen was sequenced by 7 laboratories using 10 different NGS workflows. The specimen was compatible with all 10 workflows and presented biochemical and bioinformatic challenges similar to those of patient specimens. Only 10 of 22 challenging variants were correctly identified by all 10 workflows, and only 3 workflows detected all 22. Many, but not all, of the sensitivity limitations were bioinformatic in nature. We conclude that Synthetic controls can provide an efficient and informative mechanism to augment studies with technically challenging variants that are difficult to obtain otherwise. Data from such specimens can facilitate inter-laboratory methodologic comparisons and can help establish standards that improve communication between clinicians and laboratories.

genetics

CrowdVariant: a crowdsourcing approach to classify copy number variants

Copy number variants (CNVs) are an important type of genetic variation and play a causal role in many diseases. However, they are also notoriously difficult to identify accurately from next-generation sequencing (NGS) data. For larger CNVs, genotyping arrays provide reasonable benchmark data, but NGS allows us to assay a far larger number of small (< 10kbp) CNVs that are poorly captured by array-based methods. The lack of high quality benchmark callsets of small-scale CNVs has limited our ability to assess and improve CNV calling algorithms for NGS data. To address this issue we developed a crowdsourcing framework, called CrowdVariant, that leverages Googles high-throughput crowdsourcing platform to create a high confidence set of copy number variants for NA24385 (NIST HG002/RM 8391), an Ashkenazim reference sample developed in partnership with the Genome In A Bottle Consortium. In a pilot study we show that crowdsourced classifications, even from non-experts, can be used to accurately assign copy number status to putative CNV calls and thereby identify a high-quality subset of these calls. We then scale our framework genome-wide to identify 1,781 high confidence CNVs, which multiple lines of evidence suggest are a substantial improvement over existing CNV callsets, and are likely to prove useful in benchmarking and improving CNV calling algorithms. Our crowdsourcing methodology may be a useful guide for other genomics applications.

genomics