Search bioRxivSearch

Biology subjects

Joan H.M. Knoll

Publications and source records attributed to Joan H.M. Knoll.

3 recordsLinked to original sources

Prioritizing variants in complete Hereditary Breast and Ovarian Cancer (HBOC) genes in patients lacking known BRCA mutations

BRCA1 and BRCA2 testing for HBOC does not identify all pathogenic variants. Sequencing of 20 complete genes in HBOC patients with uninformative test results (N=287), including non-coding and flanking sequences of ATM, BARD1, BRCA1, BRCA2, CDH1, CHEK2, EPCAM, MLH1, MRE11A, MSH2, MSH6, MUTYH, NBN, PALB2, PMS2, PTEN, RAD51B, STK11, TP53, and XRCC2, identified 38,372 unique variants. We apply information theory (IT) to predict novel functions for and prioritize non-coding variants of uncertain significance (VUS) in throughout regulatory, coding, and intronic regions based on changes in binding sites in these genesof these genes. Besides mRNA splicing, IT provides a common framework to evaluate potential affinity changes inin transcription factor (TFBSs), splicing regulatory (SRBSs), and RNA-binding protein (RBBSs) protein binding sites following mutationat mutated binding sites. We prioritized variants affecting the strengths of 10 variants affecting splice sites (4 natural, 6 cryptic), 148 SRBS, 36 TFBS, and 31 RBBS binding strength-affecting variantss. Three variants were also prioritized based on their predicted effects on mRNA secondary (2{degrees}) structure, and 17 for pseudoexon activation. Additionally, 4 frameshift, 2 in-frame deletions, and 5 stop-gain mutations were identified. When combined with pedigree information, complete gene sequence analysis can focus attention on a limited set of variants in a wide spectrum of functional mutation types for downstream functional and co-segregation analysis.

Genetics

Centromere Detection of Human Metaphase Chromosome Images using a Candidate Based Method

Accurate detection of the human metaphase chromosome centromere is an critical element of cytogenetic diagnostic techniques, including chromosome enumeration, karyotyping and radiation biodosimetry. Existing image processing methods can perform poorly in the presence of irregular boundaries, shape variations and premature sister chromatid separation, which can adversely affect centromere localization. We present a centromere detection algorithm that uses a novel profile thickness measurement technique on irregular chromosome structures defined by contour partitioning. Our algorithm generates a set of centromere candidates which are then evaluated based on a set of features derived from images of chromosomes. Our method also partitions the chromosome contour to isolate its telomere regions and then detects and corrects for sister chromatid separation. When tested with a chromosome database consisting of 1400 chromosomes collected from 40 metaphase cell images, the candidate based centromere detection algorithm was able to accurately localize 1220 centromere locations yielding a detection accuracy of 87%. We also introduce a Candidate Based Centromere Confidence (CBCC) metric which indicates an approximate confidence value of a given centromere detection and can be readily extended into other candidate related detection problems.

Genetics

A unified analytic framework for prioritization of non-coding variants of uncertain significance in heritable breast and ovarian cancer

BackgroundSequencing of both healthy and disease singletons yields many novel and low frequency variants of uncertain significance (VUS). Complete gene and genome sequencing by next generation sequencing (NGS) significantly increases the number of VUS detected. While prior studies have emphasized protein coding variants, non-coding sequence variants have also been proven to significantly contribute to high penetrance disorders, such as hereditary breast and ovarian cancer (HBOC). We present a strategy for analyzing different functional classes of non-coding variants based on information theory (IT).\n\nMethodsWe captured and enriched for coding and non-coding variants in genes known to harbor mutations that increase HBOC risk. Custom oligonucleotide baits spanning the complete coding, non-coding, and intergenic regions 10 kb up- and downstream of ATM, BRCA1, BRCA2, CDH1, CHEK2, PALB2, and TP53 were synthesized for solution hybridization enrichment. Unique and divergent repetitive sequences were sequenced in 102 high-risk patients without identified mutations in BRCA1/2. Aside from protein coding changes, IT-based sequence analysis was used to identify and prioritize pathogenic non-coding variants that occurred within sequence elements predicted to be recognized by proteins or protein complexes involved in mRNA splicing, transcription, and untranslated region (UTR) binding and structure. This approach was supplemented by in silico and laboratory analysis of UTR structure.\n\nResults15,311 unique variants were identified, of which 245 occurred in coding regions. With the unified IT-framework, 132 variants were identified and 87 functionally significant VUS were further prioritized. We also identified 4 stop-gain variants and 3 reading-frame altering exonic insertions/deletions (indels).\n\nConclusionsWe have presented a strategy for complete gene sequence analysis followed by a unified framework for interpreting non-coding variants that may affect gene expression. This approach distills large numbers of variants detected by NGS to a limited set of variants prioritized as potential deleterious changes.

Genetics