Search bioRxivSearch

Biology subjects

Ira M Hall

Publications and source records attributed to Ira M Hall.

3 recordsLinked to original sources

SVScore: An Impact Prediction Tool For Structural Variation

MotivationStructural variation (SV) is an important and diverse source of human genome variation. Over the past several years, much progress has been made in the area of SV detection, but predicting the functional impact of SVs discovered in whole genome sequencing (WGS) studies remains extremely challenging. Accurate SV impact prediction is especially important for WGS-based rare variant association studies and studies of rare disease.\n\nResultsHere we present SVScore, a computational tool for in silico SV impact prediction. SVScore aggregates existing per-base single nucleotide polymorphism pathogenicity scores across relevant genomic intervals for each SV in a manner that considers variant type, gene features, and uncertainty in breakpoint location. We show that in a Finnish cohort, the allele frequency spectrum of SVs with high impact scores is strongly skewed toward lower frequencies, suggesting that these variants are under purifying selection. We further show that SVScore identifies deleterious variants more effectively than naive alternative methods. Finally, our results indicate that high-scoring tandem duplications may be under surprisingly strong selection relative to high-scoring deletions, suggesting that duplications may be more deleterious than previously thought. In conclusion, SVScore provides pathogenicity prediction for SVs that is both informative and meaningful for understanding their functional role in disease.\n\nAvailabilitySVScore is implemented in Perl and available freely at {{http://www.github.com/lganel/SVScore}} for use under the MIT license.\n\nContactihall@wustl.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Bioinformatics

The impact of structural variation on human gene expression

Structural variants (SVs) are an important source of human genetic diversity but their contribution to traits, disease, and gene regulation remains unclear. The Genotype-Tissue Expression (GTEx) project presents an unprecedented opportunity to address this question due to the availability of deep whole genome sequencing (WGS) and multi-tissue RNA-seq data from 147 individuals. We used comprehensive methods to identify 24,157 high confidence SVs, and mapped cis expression quantitative trait loci (eQTLs) in 13 tissues via joint analysis of SVs, single nucleotide (SNV) and short insertion/deletion (indel) variants. We identified 24,801 eQTLs affecting the expression of 10,101 distinct genes. Based on haplotype structure and heritability partitioning, we estimate that SVs are the causal variant at 3.3-7.0% of eQTLs, which is nearly an order of magnitude higher than prior estimates from low coverage WGS and represents a 26- to 54-fold enrichment relative to their scarcity in the genome. Expression-altering SVs also have significantly larger effect sizes than SNVs and indels. We identified 787 putatively causal SVs predicted to directly alter gene expression, most of which (88.3%) are noncoding variants that show significant enrichment at enhancers and other regulatory elements. By evaluating linkage disequilibrium between SVs, SNVs and indels, we nominate 49 SVs as plausible causal variants at published genome-wide association study (GWAS) loci. Remarkably, 29.9% of the common SV-eQTLs are not well tagged by flanking SNVs, and we observe a notable abundance (relative to SNVs and indels) of rare, high impact SVs associated with aberrant expression of nearby genes. These results suggest that comprehensive WGS-based SV analyses will increase the power of both common and rare variant association studies.

Genomics

SpeedSeq: Ultra-fast personal genome analysis and interpretation

Comprehensive interpretation of human genome sequencing data is a challenging bioinformatic problem that typically requires weeks of analysis, with extensive hands-on expert involvement. This informatics bottleneck inflates genome sequencing costs, poses a computational burden for large-scale projects, and impedes the adoption of time-critical clinical applications such as personalized cancer profiling and newborn disease diagnosis, where the actionable timeframe can measure in hours or days. We developed SpeedSeq, an open-source genome analysis platform that vastly reduces computing time. SpeedSeq accomplishes read alignment, duplicate removal, variant detection and functional annotation of a 50X human genome in <24 hours, even using one low-cost server. SpeedSeq offers competitive or superior performance to current methods for detecting germline and somatic single nucleotide variants (SNVs), indels, and structural variants (SVs) and includes novel functionality for SV genotyping, SV annotation, fusion gene detection, and rapid identification of actionable mutations. SpeedSeq will help bring timely genome analysis into the clinical realm.\n\nAvailability: SpeedSeq is available at https://github.com/cc2qe/speedseq.

Bioinformatics