Search bioRxivSearch

Biology subjects

Hu, S.

Publications and source records attributed to Hu, S..

23 records · Page 2Linked to original sources

Genome-wide Variants of Eurasian Facial Shape Differentiation and a prospective model of DNA based Face Prediction

It is a long standing question as to which genes define the characteristic facial features among different ethnic groups. In this study, we use Uyghurs, an ancient admixed population to query the genetic bases why Europeans and Han Chinese look different. Facial traits were analyzed based on high-dense 3D facial images; numerous biometric spaces were examined for divergent facial features between European and Han Chinese, ranging from inter-landmark distances to dense shape geometrics. Genome-wide association analyses were conducted on a discovery panel of Uyghurs. Six significant loci were identified four of which, rs1868752, rs118078182, rs60159418 at or near UBASH3B, COL23A1, PCDH7 and rs17868256 were replicated in independent cohorts of Uyghurs or Southern Han Chinese. A prospective model was also developed to predict 3D faces based on top GWAS signals, and tested in hypothetic forensic scenarios.

genetics

CAM: A Quality Control Pipeline For MNase-Seq Data

Nucleosome organization affects the accessibility of cis-elements to trans-acting factors. Micrococcal nuclease digestion followed by high-throughput sequencing (MNase-seq) is the most popular technology used to profile nucleosome organization on a genome-wide scale. Evaluating the data quality of MNase-seq data remains challenging, especially in mammalian. There is a strong need for a convenient and comprehensive approach to obtain dedicated quality control (QC) for MNase-seq data analysis. Here we developed CAM, which is a comprehensive QC pipeline for MNase-seq data. The CAM pipeline provides multiple informative QC measurements and nucleosome organization profiles on different potentially functional regions for given MNase-seq data. CAM also includes 268 historical MNase-seq datasets from human and mouse as a reference atlas for unbiased assessment. CAM is freely available at: http://www.tongji.edu.cn/~zhanglab/CAM

bioinformatics

Dr.seq2: A Quality Control And Analysis Pipeline For Parallel Single Cell Transcriptome And Epigenome Data

An increasing number of single cell transcriptome and epigenome technologies, including single cell ATAC-seq (scATAC-seq), have been recently developed as powerful tools to analyze the features of many individual cells simultaneously. However, the methods and software were designed for one certain data type and only for single cell transcriptome data. A systematic approach for epigenome data and multiple types of transcriptome data is needed to control data quality and to perform cell-to-cell heterogeneity analysis on these ultra-high-dimensional transcriptome and epigenome datasets. Here we developed Dr.seq2, a Quality Control (QC) and analysis pipeline for multiple types of single cell transcriptome and epigenome data, including scATAC-seq and Drop-ChIP data. Application of this pipeline provides four groups of QC measurements and different analyses, including cell heterogeneity analysis. Dr.seq2 produced reliable results on published single cell transcriptome and epigenome datasets. Overall, Dr.seq2 is a systematic and comprehensive QC and analysis pipeline designed for parallel single cell transcriptome and epigenome data. Dr.seq2 is freely available at: http://www.tongji.edu.cn/~zhanglab/drseq2/ and https://github.com/ChengchenZhao/DrSeq2.

bioinformatics

Worldwide Population Structure Of Escherichia coli Reveals Two Major Subspecies

Recombination is one of the most important mechanisms of prokaryotic species evolution but its exact roles are still in debate. Here we try to infer genome-wide recombination events within a species uti-lizing a dataset of 104 complete genomes of Escherichia coli from diverse origins, among which 45 from world-wide animal-hosts are in-house sequenced using SMRT (single-molecular real time) technology.Two major clades are identified based on evidences of ecological and physiological characteristics, as well as distinct genomic features implying scarce inter-clade genetic exchange. By comparing the synteny of identical fragments genome-widely searched for each genome pair, we achieve a fine-scale map of re-combination within the population. The recombination is rather extensive within clade, which is able to break linkages between genes but does not interrupt core genome framework and primary metabolic port-folios possibly due to natural selection for physiological compatibility and ecological fitness. Meanwhile,the recombination between clades declines drastically as the phylogenetic distance increases, generally 10-fold reduced than those of the intra-clade, which establishes genetic barrier between clades. These empirical data of recombination suggest its critical role in the early stage of speciation, where recombina-tion rate differs according to phylogentic distance. The extensive intra-clade recombination coheres sister strains into a quasi-sexual group and optimizes genes or alleles to streamline physiological activities,whereas shapely declined inter-clade recombination split the population into clades adaptive to divergent ecological niches.\n\nSignificance StatementRoles of recombination in species evolution have been debated for decades due to difficulties in inferring recombination events during the early stage of speciation, especially when recombination is always complicated by frequent gene transfer events of bacterial genomes. Based on 104 high-quality complete E. coli genomes, we infer gene-centric dynamics of recombination in the formation of two E. coli clades or subpopulations, and recombination is found to be rather intensive in a within-clade fashion, which forces them to be quasi-sexual. The recombination events can be mapped among individual genomes in the context of genes and their variations; decreased between-clade and increased intra-claderecombination engender a genetic barrier that further encourages clade-specific secondary metabolic portfolios for better environmental adaptation. Recombination is thus a major force that accelerates bacterial evolution to fit ecological diversity.

microbiology

A toolbox of immunoprecipitation-grade monoclonal antibodies against human transcription factors.

A key component to overcoming the reproducibility crisis in biomedical research is the development of readily available, rigorously validated and renewable protein affinity reagents. As part of the NIH Protein Capture Reagents Program (PCRP), we have generated a collection of 1406 highly validated, immunoprecipitation (IP) and/or immunoblotting (IB) grade, mouse monoclonal antibodies (mAbs) to 736 human transcription factors. We used HuProt human protein microarrays to identify mAbs that recognize their cognate targets with exceptional specificity. Using an integrated production and validation pipeline, we validated these mAbs in multiple experimental applications, and have distributed them to the Developmental Studies Hybridoma Bank (DSHB) and several commercial suppliers. This study allowed us to perform a meta-analysis that identified critical variables that contribute to the generation of high quality mAbs. We find that using full-length antigens for immunization, in combination with HuProt analysis, provides the highest overall success rates. The efficiencies built into this pipeline ensure substantial cost savings compared to current standard practices.

biochemistry