Search bioRxivSearch

Biology subjects

Haeussler, M.

Publications and source records attributed to Haeussler, M..

7 recordsLinked to original sources

The UCSC Repeat Browser allows discovery and visualization of evolutionary conflict across repeat families

BackgroundNearly half the human genome consists of repeat elements, most of which are retrotransposons, and many of these sequences play important biological roles. However repeat elements pose several unique challenges to current bioinformatic analyses and visualization tools, as short repeat sequences can map to multiple genomic loci resulting in their misclassification and misinterpretation. In fact, sequence data mapping to repeat elements are often discarded from analysis pipelines. Therefore, there is a continued need for standardized tools and techniques to interpret genomic data of repeats. ResultsWe present the UCSC Repeat Browser, which consists of a complete set of human repeat reference sequences derived from the gold standard repeat database RepeatMasker. The UCSC Repeat Browser contains mapped annotations from the human genome to these references, and presents all of them as a comprehensive interface to facilitate work with repetitive elements. Furthermore, it provides processed tracks of multiple publicly available datasets of biological interest to the repeat community, including ChIP-SEQ datasets for KRAB Zinc Finger Proteins (KZNFs) - a family of proteins known to bind and repress certain classes of repeats. Here we show how the UCSC Repeat Browser in combination with these datasets, as well as RepeatMasker annotations in several non-human primates, can be used to trace the independent trajectories of species-specific evolutionary conflicts. ConclusionsThe UCSC Repeat Browser allows easy and intuitive visualization of genomic data on consensus repeat elements, circumventing the problem of multi-mapping, in which sequencing reads of repeat elements map to multiple locations on the human genome. By developing a reference consensus, multiple datasets and annotation tracks can easily be overlaid to reveal complex evolutionary histories of repeats in a single interactive window. Specifically, we use this approach to retrace the history of several primate specific LINE-1 families across apes, and discover several species-specific routes of evolution that correlate with the emergence and binding of KZNFs.

genomics

KRAB Zinc Finger Proteins coordinate across evolutionary time scales to battle retroelements

KRAB Zinc Finger Proteins (KZNFs) are the largest and fastest evolving family of human transcription factors1,2. The evolution of this protein family is closely linked to the tempo of retrotransposable element (RTE) invasions, with specific KZNF family members demonstrated to transcriptionally repress specific families of RTEs3,4. The competing selective pressures between RTEs and the KZNFs results in evolutionary arms races whereby KZNFs evolve to recognize RTEs, while RTEs evolve to escape KZNF recognition5. Evolutionary analyses of the primate-specific RTE family L1PA and two of its KZNF binders, ZNF93 and ZNF649, reveal specific nucleotide and amino changes consistent with an arms race scenario. Our results suggest a model whereby ZNF649 and ZNF93 worked together to target independent motifs within the L1PA RTE lineage. L1PA elements eventually escaped the concerted action of this KZNF \"team\" over [~]30 million years through two distinct mechanisms: a slow accumulation of point mutations in the ZNF649 binding site and a rapid, massive deletion of the entire ZNF93 binding site.

evolutionary biology

Structurally conserved primate lncRNAs are transiently expressed during human cortical differentiation and influence cell type specific genes

The cerebral cortex has expanded in size and complexity in primates, yet the underlying molecular mechanisms are obscure. We generated cortical organoids from human, chimpanzee, orangutan, and rhesus pluripotent stem cells and sequenced their transcriptomes at weekly time points for comparative analysis. We used transcript structure and expression conservation to discover thousands of expressed long non-coding RNAs (lncRNAs). Of 2,975 human, multi-exonic lncRNAs, 2,143 were structurally conserved to chimpanzee, 1,731 to orangutan, and 1,290 to rhesus. 386 human lncRNAs were transiently expressed (TrEx) and a similar expression pattern was often observed in great apes (46%) and rhesus (31%). Many TrEx lncRNAs were associated with neuroepithelium, radial glia, or Cajal-Retzius cells by single cell RNA-sequencing. 3/8 tested by ectopic expression showed [≥]2-fold effects on neural genes. This rich resource of primate expression data in early cortical development provides a framework for identifying new, potentially functional lncRNAs.

neuroscience

Human-specific NOTCH-like genes in a region linked to neurodevelopmental disorders affect cortical neurogenesis

Genetic changes causing dramatic brain size expansion in human evolution have remained elusive. Notch signaling is essential for radial glia stem cell proliferation and a determinant of neuronal number in the mammalian cortex. We find three paralogs of human-specific NOTCH2NL are highly expressed in radial glia cells. Functional analysis reveals different alleles of NOTCH2NL have varying potencies to enhance Notch signaling by interacting directly with NOTCH receptors. Consistent with a role in Notch signaling, NOTCH2NL ectopic expression delays differentiation of neuronal progenitors, while deletion accelerates differentiation. NOTCH2NL genes provide the breakpoints in typical cases of 1q21.1 distal deletion/duplication syndrome, where duplications are associated with macrocephaly and autism, and deletions with microcephaly and schizophrenia. Thus, the emergence of hominin-specific NOTCH2NL genes may have contributed to the rapid evolution of the larger hominin neocortex accompanied by loss of genomic stability at the 1q21. 1 locus and a resulting recurrent neurodevelopmental disorder.

neuroscience

HNRNPA1 promotes recognition of splice site decoys by U2AF2 in vivo

Alternative pre-mRNA splicing plays a major role in expanding the transcript output of human genes. This process is regulated, in part, by the interplay of trans-acting RNA binding proteins (RBPs) with myriad cis-regulatory elements scattered throughout pre-mRNAs. These molecular recognition events are critical for defining the protein coding sequences (exons) within pre-mRNAs and directing spliceosome assembly on non-coding regions (introns). One of the earliest events in this process is recognition of the 3 splice site by U2 small nuclear RNA auxiliary factor 2 (U2AF2). Splicing regulators, such as the heterogeneous nuclear ribonucleoprotein A1 (HNRNPA1), influence spliceosome assembly both in vitro and in vivo, but their mechanisms of action remain poorly described on a global scale. HNRNPA1 also promotes proof reading of 3ss sequences though a direct interaction with the U2AF heterodimer. To determine how HNRNPA1 regulates U2AF-RNA interactions in vivo, we analyzed U2AF2 RNA binding specificity using individual-nucleotide resolution crosslinking immunoprecipitation (iCLIP) in control- and HNRNPA1 over-expression cells. We observed changes in the distribution of U2AF2 crosslinking sites relative to the 3 splice sites of alternative cassette exons but not constitutive exons upon HNRNPA1 over-expression. A subset of these events shows a concomitant increase of U2AF2 crosslinking at distal intronic regions, suggesting a shift of U2AF2 to \"decoy\" binding sites. Of the many non-canonical U2AF2 binding sites, Alu-derived RNA sequences represented one of the most abundant classes of HNRNPA1-dependent decoys. Splicing reporter assays demonstrated that mutation of U2AF2 decoy sites inhibited HNRNPA1-dependent exon skipping in vivo. We propose that HNRNPA1 regulates exon definition by modulating the interaction of U2AF2 with decoy or bona fide 3 splice sites.

molecular biology

AMELIE accelerates Mendelian patient diagnosis directly from the primary literature

The diagnosis of Mendelian disorders requires labor-intensive literature research. Our software system AMELIE (Automatic Mendelian Literature Evaluation) greatly automates this process. AMELIE parses hundreds of thousands of full text articles to find an underlying diagnosis to explain a patients phenotypes given the patients exome. AMELIE prioritizes patient candidate genes for their likelihood of causing the patients phenotypes. Diagnosis of singleton patients (without relatives exomes) is the most time-consuming scenario. AMELIEs gene ranking method was tested on 215 singleton Mendelian patients with a clinical diagnosis. AMELIE ranked the causal gene among the top 2 in the majority (63%) of cases. Examining AMELIEs top 10 genes, amounting to 8% of 124 candidate genes with rare functional variants per patient, results in diagnosis for 95% of cases. Strikingly, training only on gene pathogenicity knowledge from 2011 leads to identical performance compared to training on current data. An accompanying analysis web portal has launched at AMELIE.stanford.edu.

genetics

Integrated Computational Guide Design, Execution, And Analysis Of Arrayed And Pooled CRISPR Genome Editing Experiments

CRISPR genome editing experiments offer enormous potential for the evaluation of genomic loci using arrayed single guide RNAs (sgRNAs) or pooled sgRNA libraries. Numerous computational tools are available to help design sgRNAs with optimal on-target efficiency and minimal off-target potential. In addition, computational tools have been developed to analyze deep sequencing data resulting from genome editing experiments. However, these tools are typically developed in isolation and oftentimes not readily translatable into laboratory-based experiments. Here we present a protocol that describes in detail both the computational and benchtop implementation of an arrayed and/or pooled CRISPR genome editing experiment. This protocol provides instructions for sgRNA design with CRISPOR, experimental implementation, and analysis of the resulting high-throughput sequencing data with CRISPResso. This protocol allows for design and execution of arrayed and pooled CRISPR experiments in 4-5 weeks by non-experts as well as computational data analysis in 1-2 days that can be performed by both computational and non-computational biologists alike.

molecular biology