Search bioRxivSearch

Biology subjects

Urban, A. E.

Publications and source records attributed to Urban, A. E..

7 recordsLinked to original sources

Allele-specific binding of RNA-binding proteins reveals functional genetic variants in the RNA

Allele-specific protein-RNA binding is an essential aspect that may reveal functional genetic variants influencing RNA processing and gene expression phenotypes. Recently, genome-wide detection of in vivo binding sites of RNA binding proteins (RBPs) is greatly facilitated by the enhanced UV crosslinking and immunoprecipitation (eCLIP) protocol. Hundreds of eCLIP-Seq data sets were generated from HepG2 and K562 cells during the ENCODE3 phase. These data afford a valuable opportunity to examine allele-specific binding (ASB) of RBPs. To this end, we developed a new computational algorithm, called BEAPR (Binding Estimation of Allele-specific Protein-RNA interaction). In identifying statistically significant ASB sites, BEAPR takes into account UV cross-linking induced sequence propensity and technical variations between replicated experiments. Using simulated data and actual eCLIP-Seq data, we show that BEAPR largely outperforms often-used methods Chi-Squared test and Fishers Exact test. Importantly, BEAPR overcomes the inherent over-dispersion problem of the other methods. Complemented by experimental validations, we demonstrate that ASB events are significantly associated with genetic regulation of splicing and mRNA abundance, supporting the usage of this method to pinpoint functional genetic variants in post-transcriptional gene regulation. Many variants with ASB patterns of RBPs were found as genetic variants with cancer or other disease relevance. About 38% of ASB variants were in linkage disequilibrium with single nucleotide polymorphisms from genome-wide association studies. Overall, our results suggest that BEAPR is an effective method to reveal ASB patterns in eCLIP and can inform functional interpretation of disease-related genetic variants.

bioinformatics

Haplotype-phased Callithrix jacchus embryonic stem cell line for genome editing using CRISPR/Cas9

Due to anatomical and physiological similarities to humans, the common marmoset (Callithrix jacchus) is an ideal organism for the study human diseases. Researchers are currently leveraging genome-editing technologies such as CRISPR/Cas9 to genetically engineer marmosets for the in vivo biomedical modeling of human neuropsychiatric and neurodegenerative diseases. The genome characterization of these cell lines greatly reinforces these transgenic efforts. It also provides the genomic contexts required for the accurate interpretation of functional genomics data. We performed haplotype-resolved whole-genome characterization for marmoset ESC line cj367 from the Wisconsin National Primate Research Center. This is the first haplotype-resolved analysis of a marmoset genome and the first whole-genome characterization of any marmoset ESC line. We identified and phased single-nucleotide variants (SNVs) and Indels across the genome. By leveraging this haplotype information, we then compiled a list of cj367 ESC allele-specific CRISPR targeting sites. Furthermore, we demonstrated successful Cas9 Endonuclease Dead (dCas9) expression and targeted localization in cj367 as well as sustained pluripotency after dCas9 transfection by teratoma assay. Lastly, we show that these ESCs can be directly induced into functional neurons in a rapid, single-step process. Our study provides a valuable set of genomic resources for primate transgenics in this post-genome era.

genetics

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

Sequencing of the Venter/HuRef genome using various strategies for the benchmarking of genome analysis tools

We produced an extensive collection of deep re-sequencing datasets for the Venter/HuRef genome using the Illumina massively-parallel DNA sequencing platform. The original Venter genome sequence is a very-high quality phased assembly based on Sanger sequencing. Therefore, researchers developing novel computational tools for the analysis of human genome sequence variation for the dominant Illumina sequencing technology can test and hone their algorithms by making variant calls from these Venter/HuRef datasets and then immediately confirm the detected variants in the Sanger assembly, freeing them of the need for further experimental validation. This process also applies to implementing and benchmarking existing genome analysis pipelines. We prepared and sequenced 200 bp and 350 bp short-insert whole-genome sequencing libraries (sequenced to 100x and 40x genomic coverages respectively) as well as 2 kb, 5 kb, and 12 kb mate-pair libraries (49x, 122x, and 145x physical coverages respectively). Lastly, we produced a linked-read library (128x physical coverage) from which we also performed haplotype phasing.

genomics

Comprehensive, Integrated, and Phased Whole-Genome Analysis of the Primary ENCODE Cell Line K562

K562 is widely used in biomedical research. It is one of three tier-one cell lines of ENCODE and also most commonly used for large-scale CRISPR/Cas9 screens. Although its functional genomic and epigenomic characteristics have been extensively studied, its genome sequence and genomic structural features have never been comprehensively analyzed. Such information is essential for the correct interpretation and understanding of the vast troves of existing functional genomics and epigenomics data for K562. We performed and integrated deep-coverage whole-genome (short-insert), mate-pair, and linked-read sequencing as well as karyotyping and array CGH analysis to identify a wide spectrum of genome characteristics in K562: copy numbers (CN) of aneuploid chromosome segments at high-resolution, SNVs and Indels (both corrected for CN in aneuploid regions), loss of heterozygosity, mega-base-scale phased haplotypes often spanning entire chromosome arms, structural variants (SVs) including small and large-scale complex SVs and non-reference retrotransposon insertions. Many SVs were phased, assembled, and experimentally validated. We identified multiple allele-specific deletions and duplications within the tumor suppressor gene FHIT. Taking aneuploidy into account, we re-analyzed K562 RNA-seq and whole-genome bisulfite sequencing data for allele-specific expression and allele-specific DNA methylation. We also show examples of how deeper insights into regulatory complexity are gained by integrating genomic variant information and structural context with functional genomics and epigenomics data. Furthermore, using K562 haplotype information, we produced an allele-specific CRISPR targeting map. This comprehensive whole-genome analysis serves as a resource for future studies that utilize K562 as well as a framework for the analysis of other cancer genomes.

genomics

Whole-genome sequencing analysis of genomic copy number variation (CNV) using low-coverage and paired-end strategies is highly efficient and outperforms array based CNV analysis

BackgroundCNV analysis is an integral component to the study of human genomes in both research and clinical settings. Array-based CNV analysis is the current first-tier approach in clinical cytogenetics. Decreasing costs in high-throughput sequencing and cloud computing have opened doors for the development of sequencing-based CNV analysis pipelines with fast turnaround times. We carry out a systematic and quantitative comparative analysis for several low-coverage whole-genome sequencing (WGS) strategies to detect CNV in the human genome.\n\nMethodsWe compared the CNV detection capabilities of WGS strategies (short-insert, 3kb-, and 5kb-insert mate-pair) each at 1x, 3x, and 5x coverages relative to each other and to 17 currently used high-density oligonucleotide arrays. For benchmarking, we used a set of Gold Standard (GS) CNVs generated for the 1000-Genomes-Project CEU subject NA12878.\n\nResultsOverall, low-coverage WGS strategies detect drastically more GS CNVs compared to arrays and are accompanied with smaller percentages of CNV calls without validation. Furthermore, we show that WGS (at [≥]1x coverage) is able to detect all seven GS deletion-CNVs >100 kb in NA12878 whereas only one is detected by most arrays. Lastly, we show that the much larger 15 Mbp Cri-du-chat deletion can be readily detected with short-insert paired-end WGS at even just 1x coverage.\n\nConclusionsCNV analysis using low-coverage WGS is efficient and outperforms the array-based analysis that is currently used for clinical cytogenetics.

genomics

Detection of complex structural variation from paired-end sequencing data

Complex structural variants (cxSVs), e.g. inversions with flanking deletions or interspersed inverted duplications, are part of human genetic diversity but their characteristics are not well delineated. Because their structures are difficult to resolve, cxSVs have been largely excluded from genome analysis and population-scale association studies. To permit large-scale detection of cxSVs from paired-end whole-genome sequencing, we developed Automated Reconstruction of Complex Variants (ARC-SV) using a novel probabilistic algorithm and a machine learning approach that leverages the new Human Pangenome Reference Consortium diploid assemblies. Using ARC-SV, we resolved, across 4,262 human genomes spanning all continental super-populations, 8,493 cxSVs belonging to 12 subclasses. Some cxSVs with population-specific signatures are shared with Neanderthals. Overall cxSVs are significantly enriched in regions prone to recombination and germline de novo mutations. Many cxSVs mark phenotypic hotspots (each significantly associated with [≥] 20 traits) identified in genome-wide association studies (GWAS), and 46.4% of all significant GWAS-SNPs catalogued to date reside within {+/-}125 kb of at least one cxSV locus. Common SNPs near cxSVs show significant trait heritability enrichment. Genomic regions affected by cxSVs are enriched for bivalent chromatin states. Rare cxSVs are enriched in neural genes and loci undergoing rapid or accelerated evolution and recently evolved cis-regulatory regions for human corticogenesis. We also identified 41 fixed loci where divergence from our most recent common ancestor is via localized cxSV. Our method and analysis framework allow for the accurate, efficient, and automatic identification of cxSVs for future population-scale studies of human disease and genome biology.

bioinformatics