Search bioRxivSearch

Biology subjects

Rice, S. V.

Publications and source records attributed to Rice, S. V..

2 recordsLinked to original sources

Pediatric Cancer Variant Pathogenicity Information Exchange (PeCanPIE): A Cloud-based Platform for Curating and Classifying Germline Variants

Variant interpretation in the era of next-generation sequencing (NGS) is challenging. While many resources and guidelines are available to assist with this task, few integrated end-to-end tools exist. Here we present \"PeCanPIE\" - the Pediatric Cancer Variant Pathogenicity Information Exchange, a web- and cloud-based platform for annotation, identification, and classification of variations in known or putative disease genes. Starting from a set of variants in Variant Call Format (VCF), variants are annotated, ranked by putative pathogenicity, and presented for formal classification using a decision-support interface based on published guidelines from the American College of Medical Genetics and Genomics (ACMG). The system can accept files containing millions of variants and handle single-nucleotide variants (SNVs), simple insertions/deletions (indels), multiple-nucleotide variants (MNVs), and complex substitutions. PeCanPIE has been applied to classify variant pathogenicity in cancer predisposition genes in two large-scale investigations involving >4,000 pediatric cancer patients, and serves as a repository for the expert-reviewed results. While PeCanPIEs web-based interface was designed to be accessible to non-bioinformaticians, its back end pipelines may also be run independently on the cloud, facilitating direct integration and broader adoption. PeCanPIE is publicly available and free for research use.

bioinformatics

VCF2CNA: A tool for efficiently detecting copy-number alteration in VCF genotype data

VCF2CNA is a web interface tool for copy-number alteration (CNA) analysis of VCF and other variant file formats. We applied it to 46 adult glioblastoma and 146 pediatric neuroblastoma samples sequenced by Illumina and Complete Genomics (CGI) platforms respectively. VCF2CNA was highly consistent with a state-of-the-art algorithm using raw sequencing data (mean F1-score=0.994) in high-quality glioblastoma samples and was robust to uneven coverage introduced by library artifacts. In the neuroblastoma set, VCF2CNA identified MYCN high-level amplifications in 31 of 32 clinically validated samples compared to 15 found by CGIs HMM-based CNA model. The findings suggest that VCF2CNA is an accurate, efficient and platform-independent tool for CNA analyses without accessing raw sequence data.

genomics