Search bioRxivSearch

Biology subjects

Zhou, B.

Publications and source records attributed to Zhou, B..

18 recordsLinked to original sources

Quantitative Proteomic Analysis of Prostate Tissue Specimens Identifies Deregulated ProteinComplexes in Primary Prostate Cancer

Prostate cancer (PCa) is the most frequently diagnosed non-skin cancer and a leading cause of mortality among males in developed countries. However, our understanding of the global changes of protein complexes within PCa tissue specimens remains very limited, although it has been well recognized that protein complexes carry out essentially all major processes in living organisms and that their deregulation drives the pathogenesis and progression of various diseases. By coupling tandem mass tagging-synchronous precursor selection-mass spectrometry/mass spectrometry/mass spectrometry (TMT-SPSMS3) with differential expression and co-regulation analyses, the present study compared the differences between protein complexes in normal prostate, low-grade PCa, and high-grade PCa tissue specimens. Globally, a large downregulated putative protein-protein interaction (PPI) network was detected in both low-grade and high-grade PCa, yet a large upregulated putative PPI network was only detected in high-grade but not low-grade PCa, compared with normal controls. To identify specific protein complexes that are deregulated in PCa, quantified proteins were mapped to protein complexes in CORUM, a collection of experimentally verified mammalian protein complexes. Differential expression analysis suggested that mitochondrial ribosomes and the fibrillin-associated protein complex were significantly overexpressed, whereas the ITGA6-ITGB4-Laminin10/12 and the P2X7 receptor signaling complexes were significantly downregulated, in PCa compared with normal prostate. Moreover, differential co-regulation analysis indicated that the assembly levels of some nuclear protein complexes involved in RNA synthesis and processing were significantly increased in low-grade PCa, and those of mitochondrial complex I and its subcomplexes were significantly increased in high-grade PCa, compared with normal prostate. In summary, the study represents the first global and quantitative comparison of protein complexes in prostate tissue specimens. It is expected to enhance our understanding of the molecular mechanisms underlying PCa development and progression in human patients, as well as lead to the discovery of novel biomarkers and therapeutic targets for precision management of PCa.

cancer biology

A Genome-Wide Association Study Identifies SNP Markers for Virulence in Magnaporthe oryzae Isolates from Sub-Saharan Africa

The fungal phytopathogen Magnaporthe oryzae causes blast disease in cereals such as rice and finger millet worldwide. In this study, we assessed genetic diversity of 160 isolates from nine sub-Saharan Africa (SSA) and other principal rice producing countries and conducted a genome-wide association study (GWAS) to identify the genomic regions associated with virulence of M. oryzae. GBS of isolates provided a large and high-quality 617K single nucleotide polymorphism (SNP) dataset. Disease ratings for each isolate was obtained by inoculating them onto differential lines and locally-adapted rice cultivars. Genome-wide association studies were conducted using the GBS dataset and sixteen disease rating datasets. Principal Component Analysis (PCA) was used an alternative to population structure analysis for studying population stratification from genotypic data. A significant association between disease phenotype and 528 SNPs was observed in six GWA analyses. Homology of sequences encompassing the significant SNPs was determined to predict gene identities and functions. Seventeen genes recurred in six GWA analyses, suggesting a strong association with virulence. Here, the putative genes/genomic regions associated with the significant SNPs are presented.

genetics

Allele-specific binding of RNA-binding proteins reveals functional genetic variants in the RNA

Allele-specific protein-RNA binding is an essential aspect that may reveal functional genetic variants influencing RNA processing and gene expression phenotypes. Recently, genome-wide detection of in vivo binding sites of RNA binding proteins (RBPs) is greatly facilitated by the enhanced UV crosslinking and immunoprecipitation (eCLIP) protocol. Hundreds of eCLIP-Seq data sets were generated from HepG2 and K562 cells during the ENCODE3 phase. These data afford a valuable opportunity to examine allele-specific binding (ASB) of RBPs. To this end, we developed a new computational algorithm, called BEAPR (Binding Estimation of Allele-specific Protein-RNA interaction). In identifying statistically significant ASB sites, BEAPR takes into account UV cross-linking induced sequence propensity and technical variations between replicated experiments. Using simulated data and actual eCLIP-Seq data, we show that BEAPR largely outperforms often-used methods Chi-Squared test and Fishers Exact test. Importantly, BEAPR overcomes the inherent over-dispersion problem of the other methods. Complemented by experimental validations, we demonstrate that ASB events are significantly associated with genetic regulation of splicing and mRNA abundance, supporting the usage of this method to pinpoint functional genetic variants in post-transcriptional gene regulation. Many variants with ASB patterns of RBPs were found as genetic variants with cancer or other disease relevance. About 38% of ASB variants were in linkage disequilibrium with single nucleotide polymorphisms from genome-wide association studies. Overall, our results suggest that BEAPR is an effective method to reveal ASB patterns in eCLIP and can inform functional interpretation of disease-related genetic variants.

bioinformatics

Haplotype-phased Callithrix jacchus embryonic stem cell line for genome editing using CRISPR/Cas9

Due to anatomical and physiological similarities to humans, the common marmoset (Callithrix jacchus) is an ideal organism for the study human diseases. Researchers are currently leveraging genome-editing technologies such as CRISPR/Cas9 to genetically engineer marmosets for the in vivo biomedical modeling of human neuropsychiatric and neurodegenerative diseases. The genome characterization of these cell lines greatly reinforces these transgenic efforts. It also provides the genomic contexts required for the accurate interpretation of functional genomics data. We performed haplotype-resolved whole-genome characterization for marmoset ESC line cj367 from the Wisconsin National Primate Research Center. This is the first haplotype-resolved analysis of a marmoset genome and the first whole-genome characterization of any marmoset ESC line. We identified and phased single-nucleotide variants (SNVs) and Indels across the genome. By leveraging this haplotype information, we then compiled a list of cj367 ESC allele-specific CRISPR targeting sites. Furthermore, we demonstrated successful Cas9 Endonuclease Dead (dCas9) expression and targeted localization in cj367 as well as sustained pluripotency after dCas9 transfection by teratoma assay. Lastly, we show that these ESCs can be directly induced into functional neurons in a rapid, single-step process. Our study provides a valuable set of genomic resources for primate transgenics in this post-genome era.

genetics

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

Regulatory T-cells are required for neonatal heart regeneration

Previous work has elegantly demonstrated that, unlike adult mammalian heart, the neonatal heart is able to regenerate after injury from postnatal day (P) 1 to 7. Recently, macrophages are found to be required in the repair process as depletion of which abolishes endogenous regenerative capability of the neonatal heart. Nevertheless, whether innate immunity alone is sufficient for neonatal heart regeneration is obscure. Here, we investigate a hitherto novel role of FOXP3+ regulatory T-cells (Treg) in neonatal heart regeneration. Unlike their wild type counterparts, NOD/SCID mice that are deficient for T-cells but innate immune cells including macrophages fail to regenerate their injured heart as early as P3. In wild type mice, both conventional CD4+ T-cells and Treg are recruited to cardiac muscle within the first week after injury. Treatment with the lytic anti-CD4 antibody that specifically depletes conventional CD4+ T-cells leads to reduced cardiac fibrosis; while treatment with the lytic anti-CD25 antibody that specifically depletes CD4+CD25hiFOXP3+ Treg contributes to increased fibrosis of the neonatal heart after injury. Moreover, adoptive transfer of Treg to NOD/SCID mice results in mitigated fibrosis and increased proliferation and function of cardiac muscle of the neonatal heart after injury. Mechanistically, single cell transcriptomic profiling reveals that Treg are a source of chemokines and cytokines that attract monocytes and macrophages previously known to drive neonatal heart regeneration. Furthermore, Treg directly promote proliferation of both mouse and human cardiomyocytes in a paracrine manner. Our findings uncover an unappreciated mechanism in neonatal heart regeneration; and offer new avenues for developing novel therapeutics targeting Treg-mediated heart regeneration.

developmental biology

Spatiotemporal Gene Coexpression and Regulation in Mouse Cardiomyocytes of Early Cardiac Morphogenesis

Cardiac looping is an early morphogenic process critical for the formation of four-chambered mammalian hearts. To study the roles of signaling pathways, transcription factors (TFs) and genetic networks in the process, we constructed gene co-expression networks and identified gene modules highly activated in individual cardiomyocytes (CMs) at multiple anatomical regions and developmental stages. Function analyses of the module genes uncovered major pathways important for spatiotemporal CM differentiation. Interestingly, about half of the pathways were highly active in cardiomyocytes at outflow tract (OFT) and atrioventricular canal (AVC), including many well-known signaling pathways for cardiac development and several newly identified ones. Most of the OFT-AVC pathways were predicted to be regulated by 6 6 transcription factors (TFs) actively expressed at the OFT-AVC locations, with the prediction supported by motif enrichment analysis of the TF targets, including 10 TFs that have not been previously associated with cardiac development, e.g., Etv5, Rbpms, and Baz2b. Finally, our study showed that the OFT-AVC TF targets were significantly enriched with genes associated with mouse heart developmental abnormalities and human congenital heart defects.

genetics

Integrating Hi-C and FISH data for modeling 3D organizations of chromosomes

The new advances in various experimental techniques that provide complementary in-formation about the spatial conformations of chromosomes have inspired researchers to develop computational methods to fully exploit the merits of individual data sources and combine them to improve the modeling of chromosome structure. In this paper, we propose GEM-FISH, a first method for reconstructing the 3D models of chromosomes through systematically integrating both Hi-C and FISH data with the prior biophysical knowledge of a polymer model. Comprehensive tests on a set of chromosomes for which both Hi-C and FISH data were available have demonstrated that GEM-FISH can reconstruct the 3D models of chromosomes with more accurate spatial organizations of TADs and compartments than using only Hi-C data. In addition, GEM-FISH can accurately capture the spatial proximity of loop loci and the colocalization of loci from the same sub-compartments. Moreover, our reconstructed 3D models of chromosomes revealed novel patterns of spatial distributions of super-enhancers which can provide useful insights into understanding the functional roles of these super-enhancers in gene regulation. All these results demonstrated that, through integrating both Hi-C and FISH data into a unified framework, GEM-FISH can provide a better tool for modeling the 3D organizations of chromosomes than using the Hi-C data alone.

bioinformatics

Sequencing of the Venter/HuRef genome using various strategies for the benchmarking of genome analysis tools

We produced an extensive collection of deep re-sequencing datasets for the Venter/HuRef genome using the Illumina massively-parallel DNA sequencing platform. The original Venter genome sequence is a very-high quality phased assembly based on Sanger sequencing. Therefore, researchers developing novel computational tools for the analysis of human genome sequence variation for the dominant Illumina sequencing technology can test and hone their algorithms by making variant calls from these Venter/HuRef datasets and then immediately confirm the detected variants in the Sanger assembly, freeing them of the need for further experimental validation. This process also applies to implementing and benchmarking existing genome analysis pipelines. We prepared and sequenced 200 bp and 350 bp short-insert whole-genome sequencing libraries (sequenced to 100x and 40x genomic coverages respectively) as well as 2 kb, 5 kb, and 12 kb mate-pair libraries (49x, 122x, and 145x physical coverages respectively). Lastly, we produced a linked-read library (128x physical coverage) from which we also performed haplotype phasing.

genomics

The fruitENCODE project sheds light on the genetic and epigenetic basis of convergent evolution of climacteric fruit ripening

Fleshy fruit evolved independently multiple times during angiosperm history. Many climacteric fruits utilize the hormone ethylene to regulate ripening. The fruitENCODE project shows there are multiple evolutionary origins of the regulatory circuits that govern climacteric fruit ripening. Eudicot climacteric fruits with recent whole-genome duplications (WGDs) evolved their ripening regulatory systems using the duplicated floral identity genes, while others without WGD utilised carpel senescence genes. The monocot banana uses both leaf senescence and duplicated floral-identity genes, forming two interconnected regulatory circuits. H3K27me3 plays a conserved role in restricting the expression of key ripening regulators and their direct orthologs in both the ancestral dry fruit and non-climacteric fleshy fruit species. Our findings suggest that evolution of climacteric ripening was constrained by limited availability of signalling molecules and genetic and epigenetic materials, and WGD provided new resources for plants to circumvent this limit. Understanding these different ripening mechanisms makes it possible to design tailor-made ripening traits to improve quality, yield and minimize postharvest losses.\n\nOne Sentence SummaryThe fruitENCODE project discovered three evolutionary origins of the regulatory circuits that govern climacteric fruit ripening.

genomics

Comprehensive, Integrated, and Phased Whole-Genome Analysis of the Primary ENCODE Cell Line K562

K562 is widely used in biomedical research. It is one of three tier-one cell lines of ENCODE and also most commonly used for large-scale CRISPR/Cas9 screens. Although its functional genomic and epigenomic characteristics have been extensively studied, its genome sequence and genomic structural features have never been comprehensively analyzed. Such information is essential for the correct interpretation and understanding of the vast troves of existing functional genomics and epigenomics data for K562. We performed and integrated deep-coverage whole-genome (short-insert), mate-pair, and linked-read sequencing as well as karyotyping and array CGH analysis to identify a wide spectrum of genome characteristics in K562: copy numbers (CN) of aneuploid chromosome segments at high-resolution, SNVs and Indels (both corrected for CN in aneuploid regions), loss of heterozygosity, mega-base-scale phased haplotypes often spanning entire chromosome arms, structural variants (SVs) including small and large-scale complex SVs and non-reference retrotransposon insertions. Many SVs were phased, assembled, and experimentally validated. We identified multiple allele-specific deletions and duplications within the tumor suppressor gene FHIT. Taking aneuploidy into account, we re-analyzed K562 RNA-seq and whole-genome bisulfite sequencing data for allele-specific expression and allele-specific DNA methylation. We also show examples of how deeper insights into regulatory complexity are gained by integrating genomic variant information and structural context with functional genomics and epigenomics data. Furthermore, using K562 haplotype information, we produced an allele-specific CRISPR targeting map. This comprehensive whole-genome analysis serves as a resource for future studies that utilize K562 as well as a framework for the analysis of other cancer genomes.

genomics

Whole-genome sequencing analysis of genomic copy number variation (CNV) using low-coverage and paired-end strategies is highly efficient and outperforms array based CNV analysis

BackgroundCNV analysis is an integral component to the study of human genomes in both research and clinical settings. Array-based CNV analysis is the current first-tier approach in clinical cytogenetics. Decreasing costs in high-throughput sequencing and cloud computing have opened doors for the development of sequencing-based CNV analysis pipelines with fast turnaround times. We carry out a systematic and quantitative comparative analysis for several low-coverage whole-genome sequencing (WGS) strategies to detect CNV in the human genome.\n\nMethodsWe compared the CNV detection capabilities of WGS strategies (short-insert, 3kb-, and 5kb-insert mate-pair) each at 1x, 3x, and 5x coverages relative to each other and to 17 currently used high-density oligonucleotide arrays. For benchmarking, we used a set of Gold Standard (GS) CNVs generated for the 1000-Genomes-Project CEU subject NA12878.\n\nResultsOverall, low-coverage WGS strategies detect drastically more GS CNVs compared to arrays and are accompanied with smaller percentages of CNV calls without validation. Furthermore, we show that WGS (at [≥]1x coverage) is able to detect all seven GS deletion-CNVs >100 kb in NA12878 whereas only one is detected by most arrays. Lastly, we show that the much larger 15 Mbp Cri-du-chat deletion can be readily detected with short-insert paired-end WGS at even just 1x coverage.\n\nConclusionsCNV analysis using low-coverage WGS is efficient and outperforms the array-based analysis that is currently used for clinical cytogenetics.

genomics

Detection of complex structural variation from paired-end sequencing data

Complex structural variants (cxSVs), e.g. inversions with flanking deletions or interspersed inverted duplications, are part of human genetic diversity but their characteristics are not well delineated. Because their structures are difficult to resolve, cxSVs have been largely excluded from genome analysis and population-scale association studies. To permit large-scale detection of cxSVs from paired-end whole-genome sequencing, we developed Automated Reconstruction of Complex Variants (ARC-SV) using a novel probabilistic algorithm and a machine learning approach that leverages the new Human Pangenome Reference Consortium diploid assemblies. Using ARC-SV, we resolved, across 4,262 human genomes spanning all continental super-populations, 8,493 cxSVs belonging to 12 subclasses. Some cxSVs with population-specific signatures are shared with Neanderthals. Overall cxSVs are significantly enriched in regions prone to recombination and germline de novo mutations. Many cxSVs mark phenotypic hotspots (each significantly associated with [≥] 20 traits) identified in genome-wide association studies (GWAS), and 46.4% of all significant GWAS-SNPs catalogued to date reside within {+/-}125 kb of at least one cxSV locus. Common SNPs near cxSVs show significant trait heritability enrichment. Genomic regions affected by cxSVs are enriched for bivalent chromatin states. Rare cxSVs are enriched in neural genes and loci undergoing rapid or accelerated evolution and recently evolved cis-regulatory regions for human corticogenesis. We also identified 41 fixed loci where divergence from our most recent common ancestor is via localized cxSV. Our method and analysis framework allow for the accurate, efficient, and automatic identification of cxSVs for future population-scale studies of human disease and genome biology.

bioinformatics

A Large-Scale Binding and Functional Map of Human RNA Binding Proteins

Genomes encompass all the information necessary to specify the development and function of an organism. In addition to genes, genomes also contain a myriad of functional elements that control various steps in gene expression. A major class of these elements function only when transcribed into RNA as they serve as the binding sites for RNA binding proteins (RBPs), which act to control post-transcriptional processes including splicing, cleavage and polyadenylation, RNA editing, RNA localization, stability, and translation. Despite the importance of these functional RNA elements encoded in the genome, they have been much less studied than genes and DNA elements. Here, we describe the mapping and characterization of RNA elements recognized by a large collection of human RBPs in K562 and HepG2 cells. These data expand the catalog of functional elements encoded in the human genome by addition of a large set of elements that function at the RNA level through interaction with RBPs.\n\nHighlightsO_LI223 eCLIP datasets for 150 RBPs reveal a wide variety of in vivo RNA target classes.\nC_LIO_LI472 knockdown/RNA-seq profiles of 263 RBPs reveal factor-responsive targets and integration with eCLIP indicates RNA expression and splicing regulatory patterns.\nC_LIO_LI78 RNA Bind-N-Seq profiles of in vitro binding motifs reveal links between in vitro and in vivo binding and indicate that eCLIP peaks that contain in vitro motifs are more strongly associated with regulation.\nC_LIO_LI274 maps of RBP subcellular localization by immunofluorescence indicate widespread organelle-specific RNA processing regulation.\nC_LIO_LI63 ChIP-seq profiles of DNA association suggest broad interconnectivity between chromatin association and RNA processing.\nC_LI

genomics

Chance, long tails, and inference: a non-Gaussian, Bayesian theory of vocal learning in songbirds

Traditional theories of sensorimotor learning posit that animals use sensory error signals to find the optimal motor command in the face of Gaussian sensory and motor noise. However, most such theories cannot explain common behavioral observations, for example that smaller sensory errors are more readily corrected than larger errors and that large abrupt (but not gradually introduced) errors lead to weak learning. Here we propose a new theory of sensorimotor learning that explains these observations. The theory posits that the animal learns an entire probability distribution of motor commands rather than trying to arrive at a single optimal command, and that learning arises via Bayesian inference when new sensory information becomes available. We test this theory using data from a songbird, the Bengalese finch, that is adapting the pitch (fundamental frequency) of its song following perturbations of auditory feedback using miniature headphones. We observe the distribution of the sung pitches to have long, non-Gaussian tails, which, within our theory, explains the observed dynamics of learning. Further, the theory makes surprising predictions about the dynamics of the shape of the pitch distribution, which we confirm experimentally.

animal behavior and cognition

Genomic variations in paired normal controls for lung adenocarcinomas

Somatic genomic mutations in lung adenocarcinomas (LUADs) have been extensively dissected, but whether the counterpart normal lung tissues that are exposed to ambient air or tobacco smoke as the tumor tissues do, harbor genomic variations, remains unclear. Here, the genome of normal lung tissues and paired tumors of 11 patients with LUAD were sequenced, the genome sequences of counterpart normal controls (CNCs) and tumor tissues of 513 patients were downloaded from TCGA database and analyzed. In the initial screening, genomic alterations were identified in the \"normal\" lung tissues and verified by Sanger capillary sequencing. In CNCs of TCGA datasets, a mean of 0.2721 exonic variations/Mb and 5.2885 altered genes per sample were uncovered. The C:G[->]T:A transitions, a signature of tobacco carcinogen N-methyl-N-nitro-N-nitrosoguanidine, were the predominant nucleotide changes in CNCs. 16 genes had a variant rate of more than 2%, and CNC variations in MUC5B, ZXDB, PLIN4, CCDC144NL, CNTNAP3B, and CCDC180 were associated with poor prognosis whereas alterations in CHD3 and KRTAP5-5 were associated with favorable clinical outcome of the patients. This study identified the genomic alterations in CNC samples of LUADs, and further highlighted the DNA damage effect of tobacco on lung epithelial cells.

cancer biology

Resistance genes in global crop breeding networks

Resistance genes are a major tool for managing crop diseases. The crop breeder networks that exchange resistance genes and deploy them in elite varieties help to determine the global landscape of resistance and epidemics, and comprise an important system for maintaining food security. These networks function as complex adaptive systems, with associated strengths and vulnerabilities, and implications for policies to support resistance gene deployment strategies. Extensions of epidemic network analysis can be used to evaluate the multilayer agricultural networks that support and influence crop breeding networks. We evaluate the general structure of crop breeding networks for cassava, potato, rice, and wheat, which illustrate a range of public and private configurations. These systems must adapt to global change in climate and land use, the emergence of new diseases, and disruptive breeding technologies. Principles for maintaining system resilience can be applied to global resistance gene deployment. For example, both diversity and redundancy in the roles played by individual crop breeding groups (public versus private, global versus local) may support societal goals for crop production. Another principle is management of connectivity, where enhanced connectivity among crop breeders may benefit global resistance gene deployment, but increase risks to the durability of resistance genes without effective policies regarding deployment.

genetics

Baseline mutation profiling of 1134 samples of circulating cell-free DNA and blood cells from healthy individuals

The molecular alteration in circulating cell-free DNA (cfDNA) in plasma can reflect the status of the human body in a timely manner. Hence, cfDNA has emerged as important biomarkers in clinical diagnostics, particularly in cancer. However, somatic mutations are also commonly found in healthy individuals, which extensively interfere with the diagnostic results in cancer. This study was designed to examine the background somatic mutations in white blood cells (WBC) and cfDNA for healthy controls based on the sequencing data from 1134 samples, to understand the patterns and origin of mutations detected in cfDNA. We determined the mutation frequencies in both the WBC and cfDNA groups of the samples by a panel of 50 cancer-associated genes which covered 20K nucleotide regions using ultra-deep sequencing with average depth >40000 folds. Our results showed that most of mutations in cfDNA originated from WBC. We also observed that NPM1 gene was the most frequently mutant gene in both WBC and cfDNA. Our study highlighted the importance of sequencing both cfDNA and WBC, to improve the sensitivity and accuracy for calling cancer-related mutations from circulating tumor DNA, and shielded light on developing the early cancer diagnosis by cfDNA sequencing.

bioinformatics