Search bioRxivSearch

Biology subjects

Wong, W. H.

Publications and source records attributed to Wong, W. H..

9 recordsLinked to original sources

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

DeepTACT: predicting high-resolution chromatin contacts via bootstrapping deep learning

High-resolution interactions among regulatory elements are of crucial importance for the understanding of transcriptional regulation and the interpretation of disease mechanism. Hi-C technique allows the genome-wide detection of chromatin contacts. However, unless extremely deep sequencing is performed on a very large number of input cells, current Hi-C experiments do not have high enough resolution to resolve contacts among regulatory elements. Here, we develop DeepTACT, a bootstrapping deep learning model, to integrate genome sequences and chromatin accessibility data for the prediction of chromatin contacts among regulatory elements. In tests based on promoter capture Hi-C data, DeepTACT is seen to offer improved resolution over existing methods. DeepTACT analysis also identifies a class of hub promoters, which are active across cell lines, enriched in housekeeping genes, functionally related to fundamental biological processes, and capable of reflecting cell similarity. Finally, the utility of high-resolution chromatin contact information in the study of human diseases is illustrated by the association of IFNA2 and IFNA1 to coronary artery disease via an integrative analysis of GWAS data and high-resolution contacts inferred by DeepTACT.

bioinformatics

Feedback Regulation between Initiation and Maturation Networks Orchestrates the Chromatin Dynamics of Epidermal Lineage Commitment

Tissue development results from lineage-specific transcription factors (TF) programming a dynamic chromatin landscape through progressive cell fate transitions. Here, we interrogate the epigenomic landscape during epidermal differentiation and create an inference network that ranks the coordinate effects of TF-accessible regulatory element-target gene expression triplets on lineage commitment. We discover two critical transition periods: surface ectoderm initiation and keratinocyte maturation, and identify TFAP2C and p63 as lineage initiation and maturation factors, respectively. Surprisingly, we find that TFAP2C, and not p63, is sufficient to initiate surface ectoderm differentiation, with TFAP2C-initiated progenitor cells capable of maturing into functional keratinocytes. Mechanistically, TFAP2C primes the surface ectoderm chromatin landscape and induces p63 expression and binding sites, thus allowing maturation factor p63 to positively auto-regulate its expression and close a subset of the TFAP2C-initiated early program. Our work provides a general framework to infer TF networks controlling chromatin transitions that will facilitate future regenerative medicine advances.

cell biology

Sequencing of the Venter/HuRef genome using various strategies for the benchmarking of genome analysis tools

We produced an extensive collection of deep re-sequencing datasets for the Venter/HuRef genome using the Illumina massively-parallel DNA sequencing platform. The original Venter genome sequence is a very-high quality phased assembly based on Sanger sequencing. Therefore, researchers developing novel computational tools for the analysis of human genome sequence variation for the dominant Illumina sequencing technology can test and hone their algorithms by making variant calls from these Venter/HuRef datasets and then immediately confirm the detected variants in the Sanger assembly, freeing them of the need for further experimental validation. This process also applies to implementing and benchmarking existing genome analysis pipelines. We prepared and sequenced 200 bp and 350 bp short-insert whole-genome sequencing libraries (sequenced to 100x and 40x genomic coverages respectively) as well as 2 kb, 5 kb, and 12 kb mate-pair libraries (49x, 122x, and 145x physical coverages respectively). Lastly, we produced a linked-read library (128x physical coverage) from which we also performed haplotype phasing.

genomics

Integrative analysis of single cell genomics data by coupled nonnegative matrix factorizations

When different types of functional genomics data are generated on single cells from different samples of cells from the same heterogeneous population, the clustering of cells in the different samples should be coupled. We formulate this \"coupled clustering\" problem as an optimization problem, and propose the method of coupled nonnegative matrix factorizations (coupled NMF) for its solution. The method is illustrated by the integrative analysis of single cell RNA-seq and single cell ATAC-seq data.\n\nSignificance StatementsBiological samples are often heterogeneous mixtures of different types of cells. Suppose we have two single cell data sets, each providing information on a different cellular feature and generated on a different sample from this mixture. Then, the clustering of cells in the two samples should be coupled as both clusterings are reflecting the underlying cell types in the same mixture. This \"coupled clustering\" problem is a new problem not covered by existing clustering methods. In this paper we develop an approach for its solution based the coupling of two nonnegative matrix factorizations. The method should be useful for integrative single cell genomics analysis tasks such as the joint analysis of single cell RNA-seq and single cell ATAC-seq data.

bioinformatics

Comprehensive, Integrated, and Phased Whole-Genome Analysis of the Primary ENCODE Cell Line K562

K562 is widely used in biomedical research. It is one of three tier-one cell lines of ENCODE and also most commonly used for large-scale CRISPR/Cas9 screens. Although its functional genomic and epigenomic characteristics have been extensively studied, its genome sequence and genomic structural features have never been comprehensively analyzed. Such information is essential for the correct interpretation and understanding of the vast troves of existing functional genomics and epigenomics data for K562. We performed and integrated deep-coverage whole-genome (short-insert), mate-pair, and linked-read sequencing as well as karyotyping and array CGH analysis to identify a wide spectrum of genome characteristics in K562: copy numbers (CN) of aneuploid chromosome segments at high-resolution, SNVs and Indels (both corrected for CN in aneuploid regions), loss of heterozygosity, mega-base-scale phased haplotypes often spanning entire chromosome arms, structural variants (SVs) including small and large-scale complex SVs and non-reference retrotransposon insertions. Many SVs were phased, assembled, and experimentally validated. We identified multiple allele-specific deletions and duplications within the tumor suppressor gene FHIT. Taking aneuploidy into account, we re-analyzed K562 RNA-seq and whole-genome bisulfite sequencing data for allele-specific expression and allele-specific DNA methylation. We also show examples of how deeper insights into regulatory complexity are gained by integrating genomic variant information and structural context with functional genomics and epigenomics data. Furthermore, using K562 haplotype information, we produced an allele-specific CRISPR targeting map. This comprehensive whole-genome analysis serves as a resource for future studies that utilize K562 as well as a framework for the analysis of other cancer genomes.

genomics

Detection of complex structural variation from paired-end sequencing data

Complex structural variants (cxSVs), e.g. inversions with flanking deletions or interspersed inverted duplications, are part of human genetic diversity but their characteristics are not well delineated. Because their structures are difficult to resolve, cxSVs have been largely excluded from genome analysis and population-scale association studies. To permit large-scale detection of cxSVs from paired-end whole-genome sequencing, we developed Automated Reconstruction of Complex Variants (ARC-SV) using a novel probabilistic algorithm and a machine learning approach that leverages the new Human Pangenome Reference Consortium diploid assemblies. Using ARC-SV, we resolved, across 4,262 human genomes spanning all continental super-populations, 8,493 cxSVs belonging to 12 subclasses. Some cxSVs with population-specific signatures are shared with Neanderthals. Overall cxSVs are significantly enriched in regions prone to recombination and germline de novo mutations. Many cxSVs mark phenotypic hotspots (each significantly associated with [≥] 20 traits) identified in genome-wide association studies (GWAS), and 46.4% of all significant GWAS-SNPs catalogued to date reside within {+/-}125 kb of at least one cxSV locus. Common SNPs near cxSVs show significant trait heritability enrichment. Genomic regions affected by cxSVs are enriched for bivalent chromatin states. Rare cxSVs are enriched in neural genes and loci undergoing rapid or accelerated evolution and recently evolved cis-regulatory regions for human corticogenesis. We also identified 41 fixed loci where divergence from our most recent common ancestor is via localized cxSV. Our method and analysis framework allow for the accurate, efficient, and automatic identification of cxSVs for future population-scale studies of human disease and genome biology.

bioinformatics

Unsupervised Clustering And Epigenetic Classification Of Single Cells

Characterizing epigenetic heterogeneity at the cellular level is a critical problem in the modern genomics era. Assays such as single cell ATAC-seq (scATAC-seq) offer an opportunity to interrogate cellular level epigenetic heterogeneity through patterns of variability in open chromatin. However, these assays exhibit technical variability that complicates clear classification and cell type identification in heterogeneous populations. We present scABC, an R package for the unsupervised clustering of single cell epigenetic data, to classify scATAC-seq data and discover regions of open chromatin specific to cell identity.

bioinformatics

Scalable Multi-Sample Single-Cell Data Analysis by Partition-Assisted Clustering and Multiple Alignments of Networks

Mass cytometry (CyTOF) has greatly expanded the capability of cytometry. It is now easy to generate multiple CyTOF samples in a single study, with each sample containing single-cell measurement on 50 markers for more than hundreds of thousands of cells. Current methods do not adequately address the issues concerning combining multiple samples for subpopulation discovery, and these issues can be quickly and dramatically amplified with increasing number of samples. To overcome this limitation, we developed Partition-Assisted Clustering and Multiple Alignments of Networks (PAC-MAN) for the fast automatic identification of cell populations in CyTOF data closely matching that of expert manual-discovery, and for alignments between subpopulations across samples to define dataset-level cellular states. PAC-MAN is computationally efficient, allowing the management of very large CyTOF datasets, which are increasingly common in clinical studies and cancer studies that monitor various tissue samples for each subject.\n\nAuthor SummaryRecently, the cytometry field has experienced rapid advancement in the development of mass cytometry (CyTOF). CyTOF enables a significant increase in the ability to monitor 50 or more cellular markers for millions of cells at the single-cell level. Initial studies with CyTOF focused on few samples, in which expert manual discovery of cell types were acceptable. As the technology matures, it is now feasible to collect more samples, which enables systematic studies of cell types across multiple samples. However, the statistical and computational issues surrounding multi-sample analysis have not been previously examined in detail. Furthermore, it was not clear how the data analysis could be scaled for hundreds of samples, such as those in clinical studies. In this work, we present a scalable analysis pipeline that is grounded in strong statistical foundation. Partition-Assisted Clustering (PAC) offers fast and accurate clustering and Multiple Alignments of Networks (MAN) utilizes network structures learned from each homogeneous cluster to organize the data into data-set level clusters. PAC-MAN thus enables the analysis of a large CyTOF dataset that was previously too large to be analyzed systematically; this pipeline can be extended to the analysis of similarly large or larger datasets.

bioinformatics