Search bioRxivSearch

Biology subjects

Song, G.

Publications and source records attributed to Song, G..

5 recordsLinked to original sources

Haplotype-resolved and integrated genome analysis of ENCODE cell line HepG2

The HepG2 cancer cell line is one of the most widely-used biomedical research and one of the main cell lines of ENCODE. Vast numbers of functional genomics and epigenomics datasets have been produced to characterize its biology. However, the correct interpretation such data requires an understanding of the cell lines genome sequence and genome structure. Using a variety of sequencing and analysis methods, we identified a wide spectrum of HepG2 genome characteristics: copy numbers of chromosomal segments, SNVs and Indels (corrected for aneuploidy), phased haplotypes extending to entire chromosome arms, loss of heterozygosity, retrotransposon insertions, structural variants (SVs) including complex and somatic genomic rearrangements. We also identified allele-specific expression and DNA methylation genome-wide and assembled an allele-specific CRISPR/Cas9 targeting map.\n\nSIGNIFICANCEHaplotype-resolved and comprehensive whole-genome analysis of a widely-used cell line for cancer research and ENCODE, HepG2, serves as an essential resource for unlocking complex cancer gene regulation using a genome-integrated framework and also provides genomic context for the analysis of ~1,000 functional datasets to date on ENCODE for biological discovery. We also demonstrate how deeper insights into genomic regulatory complexity are gained by adopting a genome-integrated framework.

genomics

Maize defective kernel5 is a bacterial tamB homolog required for chloroplast envelope biogenesis

Chloroplasts are of prokaryotic origin with a double membrane envelope that separates plastid metabolism from the cytosol. Envelope membrane proteins integrate the chloroplast with the cell, but the biogenesis of the envelope membrane remains elusive. We show that the maize defective kernel5 (dek5) locus is critical for plastid membrane biogenesis. Amyloplasts and chloroplasts are larger and reduced in number in dek5 with multiple ultrastructural defects. We show that dek5 encodes a protein homologous to rice SUBSTANDARD STARCH GRAIN4 (SSG4) and E.coli tamB. TamB functions in bacterial outer membrane biogenesis. The DEK5 protein is localized to the chloroplast envelope with a topology analogous to TamB. Increased levels of soluble sugars in dek5 developing endosperm and elevated osmotic pressure in mutant leaf cells suggest defective intracellular solute transport. Both proteomics and antibody-based analyses show that dek5 chloroplasts have reduced levels of chloroplast envelope transporters. Moreover, dek5 chloroplasts reduce inorganic phosphate uptake with at least an 80% reduction relative to normal chloroplasts. These data suggest that DEK5 functions in plastid envelope biogenesis to enable metabolite transport.

plant biology

Assessment and refinement of sample preparation methods for deep and quantitative plant proteome profiling

A major challenge in the field of proteomics is obtaining high quality peptides for comprehensive proteome profiling by liquid chromatography mass spectrometry for many organisms. Here we evaluate and modify a range of sample preparation methods using photosynthetically active Arabidopsis leaf tissues from several developmental timepoints. We find that inclusion of FASP-based on filter digestion improves all protein extraction methods tested. Ultimately, we show that a detergent-free urea-FASP approach enables deep and robust quantification of leaf proteomes. For example, from 4-day-old leaf tissue we profiled up to 11,690 proteins from a single sample replicate. This method should be broadly applicable to researchers working on difficult to process samples from a range of plant and non-plant organisms.\n\nAbbreviations

plant biology

Comprehensive, Integrated, and Phased Whole-Genome Analysis of the Primary ENCODE Cell Line K562

K562 is widely used in biomedical research. It is one of three tier-one cell lines of ENCODE and also most commonly used for large-scale CRISPR/Cas9 screens. Although its functional genomic and epigenomic characteristics have been extensively studied, its genome sequence and genomic structural features have never been comprehensively analyzed. Such information is essential for the correct interpretation and understanding of the vast troves of existing functional genomics and epigenomics data for K562. We performed and integrated deep-coverage whole-genome (short-insert), mate-pair, and linked-read sequencing as well as karyotyping and array CGH analysis to identify a wide spectrum of genome characteristics in K562: copy numbers (CN) of aneuploid chromosome segments at high-resolution, SNVs and Indels (both corrected for CN in aneuploid regions), loss of heterozygosity, mega-base-scale phased haplotypes often spanning entire chromosome arms, structural variants (SVs) including small and large-scale complex SVs and non-reference retrotransposon insertions. Many SVs were phased, assembled, and experimentally validated. We identified multiple allele-specific deletions and duplications within the tumor suppressor gene FHIT. Taking aneuploidy into account, we re-analyzed K562 RNA-seq and whole-genome bisulfite sequencing data for allele-specific expression and allele-specific DNA methylation. We also show examples of how deeper insights into regulatory complexity are gained by integrating genomic variant information and structural context with functional genomics and epigenomics data. Furthermore, using K562 haplotype information, we produced an allele-specific CRISPR targeting map. This comprehensive whole-genome analysis serves as a resource for future studies that utilize K562 as well as a framework for the analysis of other cancer genomes.

genomics

A toolbox of immunoprecipitation-grade monoclonal antibodies against human transcription factors.

A key component to overcoming the reproducibility crisis in biomedical research is the development of readily available, rigorously validated and renewable protein affinity reagents. As part of the NIH Protein Capture Reagents Program (PCRP), we have generated a collection of 1406 highly validated, immunoprecipitation (IP) and/or immunoblotting (IB) grade, mouse monoclonal antibodies (mAbs) to 736 human transcription factors. We used HuProt human protein microarrays to identify mAbs that recognize their cognate targets with exceptional specificity. Using an integrated production and validation pipeline, we validated these mAbs in multiple experimental applications, and have distributed them to the Developmental Studies Hybridoma Bank (DSHB) and several commercial suppliers. This study allowed us to perform a meta-analysis that identified critical variables that contribute to the generation of high quality mAbs. We find that using full-length antigens for immunization, in combination with HuProt analysis, provides the highest overall success rates. The efficiencies built into this pipeline ensure substantial cost savings compared to current standard practices.

biochemistry