Search bioRxivSearch

Biology subjects

Bejerano, G.

Publications and source records attributed to Bejerano, G..

12 recordsLinked to original sources

Components of genetic associations across 2,138 phenotypes in the UK Biobank highlight novel adipocyte biology

To characterize latent components of genetic associations, we applied truncated singular value decomposition (DeGAs) to matrices of summary statistics derived from genome-wide association analyses across 2,138 phenotypes measured in 337,199 White British individuals in the UK Biobank study. We systematically identified key components of genetic associations and the contributions of variants, genes, and phenotypes to each component. As an illustration of the utility of the approach to inform downstream experiments, we report putative loss of function variants, rs114285050 (GPR151) and rs150090666 (PDE3B), that substantially contribute to obesity-related traits, and experimentally demonstrate the role of these genes in adipocyte biology. Our approach to dissect components of genetic associations across human phenotypes will accelerate biomedical hypothesis generation by providing insights on previously unexplored latent structures.

genetics

Identification of rare-disease genes in diverse undiagnosed cases using whole blood transcriptome sequencing and large control cohorts

RNA sequencing (RNA-seq) is a complementary approach for Mendelian disease diagnosis for patients in whom exome-sequencing is not informative. For both rare neuromuscular and mitochondrial disorders, its application has improved diagnostic rates. However, the generalizability of this approach to diverse Mendelian diseases has yet to be evaluated. We sequenced whole blood RNA from 56 cases with undiagnosed rare diseases spanning 11 diverse disease categories to evaluate the general application of RNA-seq to Mendelian disease diagnosis. We developed a robust approach to compare rare disease cases to existing large sets of RNA-seq controls (N=1,594 external and N=31 family-based controls) and demonstrated the substantial impacts of gene and variant filtering strategies on disease gene identification when combined with RNA-seq. Across our cohort, we observed that RNA-seq yields a 8.5% diagnostic rate. These diagnoses included diseases where blood would not intuitively reflect evidence of disease. We identified RARS2 as an under-expression outlier containing compound heterozygous pathogenic variants for an individual exhibiting profound global developmental delay, seizures, microcephaly, hypotonia, and progressive scoliosis. We also identified a new splicing junction in KCTD7 for an individual with global developmental delay, loss of milestones, tremors and seizures. Our study provides a broad evaluation of blood RNA-seq for the diagnosis of rare disease.

genomics

ClinPhen extracts and prioritizes patientphenotypes directly from medical records to accelerate genetic disease diagnosis

PurposeSevere genetic diseases affect 7 million births per year, worldwide. Diagnosing these diseases is necessary for optimal care, but it can involve the manual evaluation of hundreds of genetic variants per case, with many variants taking an hour to evaluate. Automatic gene-ranking approaches shorten this process by reporting which of the genes containing variants are most likely to be causing the patients symptoms. To use these tools, busy clinicians must manually encode patient phenotypes, which is a cumbersome and imprecise process. With 60 million patients expected to be sequenced in the next 7 years, a fast alternative to manual phenotype extraction from the clinical notes in patients medical records will become necessary.\n\nMethodsWe introduce ClinPhen: a fast, high-accuracy tool that automatically converts the clinical notes into a prioritized list of patient symptoms using HPO terms.\n\nResultsClinPhen shows superior accuracy to existing phenotype extractors, and when paired with a gene-ranking tool it significantly improve the latters performance.\n\nConclusionCompared to manual phenotype extraction, ClinPhen saves more than 5 hours per case in Mendelian diagnosis alone. Summing over millions of forthcoming cases whose medical notes await phenotype encoding, ClinPhen makes a substantial contribution towards ending all patients diagnostic odyssey.

genomics

S-CAP extends clinical-grade pathogenicity prediction to genetic variants that affect RNA splicing

There are over 15,000 known variants that cause human inherited disease by disrupting RNA splicing. While several in silico methods such as CADD, EIGEN and LINSIGHT are commonly used to predict the pathogenicity of noncoding variants, we introduce S-CAP, a tool developed specially for splicing which is better able to effectively distinguish pathogenic splicing-relevant variants from benign variants. S-CAP is a novel splicing pathogenicity predictor that reduces the number of splicing-relevant variants of uncertain significance in patient exomes by 41%, a nearly 3-fold improvement over existing noncoding pathogenicity measures while correctly classifying known pathogenic splicing-relevant variants with a clinical-grade 95% sensitivity.

genomics

Phrank measures phenotype sets similarity to greatly improve Mendelian diagnostic disease prioritization

PurposeExome sequencing and diagnosis is beginning to spread across the medical establishment. The most time-consuming part of genome based diagnosis is the manual step of matching the potentially long list of patient candidate genes to patient phenotypes to identify the causative disease.\n\nMethodsWe introduce Phrank (for phenotype ranking), an information-theory inspired method that utilizes a Bayesian Network to prioritize candidate diseases or genes, as a stand-alone module that can be run with any underlying knowledgebase and any variant filtering scheme.\n\nResultsPhrank outperforms existing methods at ranking the causative disease or gene when applied to 169 real patient exomes with Mendelian diagnoses. Phranks greatest improvement is in disease space, where across all 169 patients it ranks only 3 diseases on average ahead of the true diagnosis, whereas Phenomizer ranks 32 diseases ahead of the causal one.\n\nConclusionUsing Phrank to rank all patient candidate genes or diseases, as they start working through a new case, will save the busy clinician much time in deriving a genetic diagnosis.

genomics

Independent erosion of conserved transcription factor binding sites points to shared hindlimb, vision, and scrotum loss in different mammals

Genetic variation in cis-regulatory elements is thought to be a major driving force in morphological and physiological change. However, identifying transcription factor binding events which code for complex traits remains a challenge, motivating novel means of detecting putatively important binding events. Using a curated set of 1,154 high-quality transcription factor motifs, we demonstrate that independently eroded binding sites are enriched for independently lost traits in three distinct pairs of placental mammals. We show that these independently eroded events pinpoint the loss of hindlimbs in dolphin and manatee, degradation of vision in naked mole-rat and star-nosed mole, and the loss of scrotum in white rhinoceros and Weddell seal. Our study exhibits a novel methodology to detect cis-regulatory mutations which help explain a portion of the molecular mechanism underlying complex trait formation and loss.\n\nAuthor SummaryEvolution has produced an astounding variety of species with incredibly diverse phenotypes. A central question in evolutionary developmental biology is how (and which) DNA evolves to encode all of these different traits. A prevailing hypothesis is that changes in regulatory DNA, short stretches of DNA which control the expression of protein-coding genes, drive important differences in trait formation between species. The basic building block of regulatory DNA is thought to be transcription factor binding sites, shortl genomic sequences which attract proteins whose central role is to control the rate of transcription. In this study, we asked whether the independent erosion of otherwise highly conserved transcription factor binding sites points to a trait shared between species which have undergone similar adaptations. We show that our method is able to point to the loss of hindlimbs in dolphin and manatee, poor vision in naked mole-rat and star-nosed mole, and loss of scrotum in Weddell seal and white rhinoceros. Overall, our study exhibits a means of detecting evolutionarily important genomic regions which help explain a portion of complex trait loss and retention.

evolutionary biology

Comment on "A genetic signature of the evolution of loss of flight in the Galapagos cormorant"

Burga et al. could not identify gene regulatory regions which might contribute to flightlessness in the Galapagos cormorant. Using a bird-specific alignment we discover 48 limb enhancers showing strong accelerated evolution in P. harrisi, including enhancers of key cilia development, hedgehog signaling, and planar cell polarity genes, such as Prickle1, Twist2 and Ndk1, extending Burga's proposed mechanism into the non-coding genome.

genomics

A sequence-based, deep learning model accurately predicts RNA splicing branchpoints

Experimental detection of RNA splicing branchpoints, the nucleotide serving as the nucleophile in the first catalytic step of splicing, is difficult. To date, annotations exist for only 16-21% of 3 splice sites in the human genome and even these limited annotations have been shown to be plagued by noise. We develop a sequence-only, deep learning based branchpoint predictor, LaBranchoR, which we conclude predicts a correct branchpoint for over 90% of 3 splice sites genome-wide. Our predicted branchpoints show large agreement with trends observed in the raw data, but analysis of conservation signatures and overlap with pathogenic variants reveal that our predicted branchpoints are generally more reliable than the raw data itself. We use our predicted branchpoints to identify a sequence element upstream of branchpoints consistent with extended U2 snRNA base pairing, show an association between weak branchpoints and alternative splicing, and explore the effects of variants on branchpoints.

bioinformatics

AMELIE accelerates Mendelian patient diagnosis directly from the primary literature

The diagnosis of Mendelian disorders requires labor-intensive literature research. Our software system AMELIE (Automatic Mendelian Literature Evaluation) greatly automates this process. AMELIE parses hundreds of thousands of full text articles to find an underlying diagnosis to explain a patients phenotypes given the patients exome. AMELIE prioritizes patient candidate genes for their likelihood of causing the patients phenotypes. Diagnosis of singleton patients (without relatives exomes) is the most time-consuming scenario. AMELIEs gene ranking method was tested on 215 singleton Mendelian patients with a clinical diagnosis. AMELIE ranked the causal gene among the top 2 in the majority (63%) of cases. Examining AMELIEs top 10 genes, amounting to 8% of 124 candidate genes with rare functional variants per patient, results in diagnosis for 95% of cases. Strikingly, training only on gene pathogenicity knowledge from 2011 leads to identical performance compared to training on current data. An accompanying analysis web portal has launched at AMELIE.stanford.edu.

genetics

A novel unbiased test for molecular convergent evolution and discoveries in echolocating, aquatic and high-altitude mammals

Distantly related species entering similar biological niches often adapt by evolving similar morphological and physiological characters. The extent to which genomic molecular convergence, and the extent to which coding mutations underlie this convergent phenotypic evolution remain unknown. Using a novel test, we ask which group of functionally coherent genes is most affected by convergent amino acid substitutions between phenotypically convergent lineages. This most affected sets reveals 75 novel coding convergences in important genes that pattern a highly adapted organ: the cochlea, skin and lung in echolocating, aquatic and high-altitude mammals, respectively. Our test explicitly requires the enriched converged term to not be simultaneously enriched for divergent mutations, and correctly dismisses relaxation-based signals, such as those produced by vision genes in subterranean mammals. This novel test can be readily applied to birds, fish, flies, worms etc., to discover more of the fascinating contribution of protein coding convergence to phenotype convergence.

evolutionary biology

Revealing the causative variant in Mendelian patient genomes without revealing patient genomes

Given the rapidly growing utility of critical health information revealed in the human genome, secure genomic computation is essential to moving forward, especially as genome sequencing becomes commonplace. We devise and implement proof-of-principle computational operations for precisely identifying causal variants in Mendelian patients using secure multiparty computation methods based on Yaos protocol. We show multiple real scenarios (small patient cohorts, trio analysis, two hospital collaboration) where the causal variant is discovered jointly, while keeping up to 99.7% of all participants most sensitive genomic information private. All similar operations performed today to diagnose such cases are done openly, keeping 0% of participants genomic information private. Our work will help usher in an era where genomes can be both utilized and truly protected.

genomics

Darwin: A Hardware-acceleration Framework for Genomic Sequence Alignment

Genomics is set to transform medicine and our understanding of life in fundamental ways. But the growth in genomics data has been overwhelming - far outpacing Moores Law. The advent of third generation sequencing technologies is providing new insights into genomic contribution to diseases with complex mutation events, but have prohibitively high computational costs. Over 1,300 CPU hours are required to align reads from a 54x coverage of the human genome to a reference (estimated using [1]), and over 15,600 CPU hours to assemble the reads de novo [2]. This paper proposes \"Darwin\" - a hardware-accelerated framework for genomic sequence alignment that, without sacrificing sensitivity, provides 125x and 15.6x speedup over the state-of-the-art software counterparts for reference-guided and de novo assembly of third generation sequencing reads, respectively. For pairwise alignment of sequences, Darwin is over 39,000x more energy-efficient than software. Darwin uses (i) a novel filtration strategy, called D-SOFT, to reduce the search space for sequence alignment at high speed, and (ii) a hardware-accelerated version of GACT, a novel algorithm to generate near-optimal alignments of arbitrarily long genomic sequences using constant memory for trace-back. Darwin is adaptable, with tunable speed and sensitivity to match emerging sequencing technologies and to meet the requirements of genomic applications beyond read assembly.

genomics