Search bioRxivSearch

Biology subjects

Zhang, L.

Publications and source records attributed to Zhang, L..

At least 55 records · Page 3Linked to original sources

Whole-exome sequencing identified rare variants associated with body length and girth in cattle

Body measurements can be used in determining body size to monitor the cattle growth and examine the response to selection. Despite efforts putting into the identification of common genetic variants, the mechanism understanding of the rare variation in complex traits about body size and growth remains limited. Here, we firstly performed GWAS study for body measurement traits in Simmental cattle, however there were no SNPs exceeding significant level associated with body measurements. To further investigate the mechanism of growth traits in beef cattle, we conducted whole exome analysis of 20 cattle with phenotypic differences on body girth and length, representing the first systematic exploration of rare variants on body measurements in cattle. By carrying out a three-phase process of the variant calling and filtering, a sum of 1158, 1151, 1267, and 1303 rare variants were identified in four phenotypic groups of two growth traits, higher/ lower body girth (BG_H and BG_L) and higher/lower body length (BL_H and BL_L) respectively. The subsequent functional enrichment analysis revealed that these rare variants distributed in 886 genes associated with collagen formation and organelle organization, indicating the importance of collagen formation and organelle organization for body size growth in cattle. The integrative network construction distinguished 62 and 66 genes with different co-expression patterns associated with higher and lower phenotypic groups of body measurements respectively, and the two sub-networks were distinct. Gene ontology and pathway annotation further showed that all shared genes in phenotypic differences participate in many biological processes related to the growth and development of the organism. Together, these findings provide a deep insight into rare genetic variants of growth traits in cattle and this will have a promising application in animal breeding.

genetics

Optical Imaging of Metabolic Dynamics in Animals

Direct visualization of metabolic dynamics in living tissues with high spatial and temporal resolution is essential to understanding many biological processes. Here we introduce a platform that combines deuterium oxide (D2O) probing with stimulated Raman scattering microscopy (DO-SRS) to image in situ metabolic activities. Enzymatic incorporation of D2O-derived deuterium into macromolecules generates carbon-deuterium (C-D) bonds, which track biosynthesis in tissues and can be imaged by SRS in situ. Within the broad vibrational spectra of C-D bonds, we discovered lipid-, protein-, and DNA-specific Raman shifts and developed spectral unmixing methods to obtain C-D signals with macromolecular selectivity. DO-SRS enabled us to probe de novo lipogenesis in animals, image protein biosynthesis without tissue bias, and simultaneously visualize lipid and protein metabolism and reveal their different dynamics. DO-SRS, being noninvasive, universally applicable, and cost-effective, can be adapted to a broad range of biological systems to study development, tissue homeostasis, aging, and tumor heterogeneity.

biophysics

BisPin and BFAST-Gap: Mapping Bisulfite-Treated Reads

BackgroundBisPin is a new multiprocess bisulfite-treated short DNA read mapper written in Python 2.7. It performs alignments using BFAST, leveraging its multithreading functionality and thorough hash-based indexing strategy. BisPin is feature rich and supports directional, nondirectional, PBAT, and hairpin construction strategies. BisPin approaches read mapping by converting the Cs to Ts and the Gs to As in both the reads and the reference genome. BisPin uses fast rescoring to disambiguate ambiguously aligned reads for a superior amount of uniquely mapped reads compared to other mappers. The performance of BisPin was evaluated on both real and simulated data in comparison to other read mappers.\n\nBFAST-Gap is a modified version of BFAST meant for Ion Torrent reads. It uses a parameterized logistic function to determine the weights of the gap open and extension penalties based on the homopolymer run length of the DNA read. This is because the Ion Torrent sequencing technology can overcall and undercall homopolymer runs. BisPin works with both BFAST-Gap and BFAST. BFAST-Gap is compatible with indexes built with BFAST. There are few mappers that specifically address Ion Torrent data. BFAST-Gap works with Illumina reads as well.\n\nResultsBisPin with BFAST consistently had a higher amount of uniquely mapped reads compared to other mappers on real data using a variety of construction strategies. Using a hairpin validation strategy, BisPin was superior using the maximum score, and it mapped 73% of reads correctly.\n\nBisPin with BFAST-Gap on Ion Torrent reads with a logistic gap open penalty function improved mapping accuracy with real and simulated data. On simulated bisulfite Ion Torrent data, the area under the curve was improved by approximately seven, and on one real data set, the uniquely mapped percent was improved by seven percent. BFAST-Gap performed better than TMAP on simulated regular Ion Torrent reads, and TMAP is designed for Ion Torrent reads. Other read mappers had worse performance.\n\nConclusionsBisPin and BFAST-Gap have consistently good accuracy with a variety of data. BisPin is feature-rich. This makes BisPin and BFAST-Gap useful additions to read mapping software.

bioinformatics

Encoding the expectation of a sensory stimulus

Most organisms possess an ability to differentiate unexpected or surprising sensory stimuli from those that are repeatedly encountered. How is this sensory computation performed? We examined this issue in the locust olfactory system. We found that odor-evoked responses in the antennal lobe (downstream to sensory neurons) systematically reduced upon repeated encounters of a temporally discontinuous stimulus. Rather than confounding information about stimulus identity and intensity, neural representations were optimized to encode equivalent stimulus-specific information with fewer spikes. Further, spontaneous activity of the antennal lobe network also changed systematically and became negatively correlated with the response elicited by the repetitive stimulus (i.e. a negative image). Notably, while response to the repetitive stimulus reduced, exposure to an unexpected/deviant cue generated undamped and even exaggerated spiking responses in several neurons. In sum, our results reveal how expectation regarding a stimulus is encoded in a neural circuit to allow response optimization and preferential filtering.

neuroscience

Single-step Enzymatic Glycoengineering for the Construction of Antibody-cell Conjugates

Employing live cells as therapeutics is a direction of future drug discovery. An easy and robust method to modify the surfaces of cells directly to incorporate novel functionalities is highly desirable. However, many current methods for cell-surface engineering interfere with cells endogenous properties. Here we report an enzymatic approach that enables the transfer of biomacromolecules, such as a full length IgG antibody, to the glycocalyx on the surfaces of live cells when the antibody is conjugated to the enzymes natural donor substrate GDP-fucose. This method is fast and biocompatible with little interference to cells endogenous functions. We applied this method to construct two antibody-cell conjugates (ACCs) using different immune cells, and the modified cells exhibited specific tumor targeting and resistance to inhibitory signals produced by tumor cells, respectively. Remarkably, Herceptin-NK-92MI conjugates exhibits enhanced activities to induce the lysis of HER2+ cancer cells both ex vivo and in a murine tumor model, indicating its potential for further development as a clinical candidate.

bioengineering

Microbiome inhibition of IRAK-4 by trimethylamine mediates metabolic and immune benefits in high-fat-diet-induced insulin resistance

The global type 2 diabetes epidemic is a major health crisis and there is a critical need for innovative strategies to fight it. Although the microbiome plays important roles in the onset of insulin resistance (IR) and low-grade inflammation, the microbial compounds regulating these phenomena remain to be discovered. Here, we reveal that the microbiome inhibits a central kinase, eliciting immune and metabolic benefits. Through a series of in vivo experiments based on choline supplementation, blocking trimethylamine (TMA) production then administering TMA, we demonstrate that TMA decouples inflammation and IR from obesity in the context of high-fat diet (HFD) feeding. Through in vitro kinome screens, we reveal TMA specifically inhibits Interleukin-1 Receptor-associated Kinase 4 (IRAK4), a central kinase integrating signals from various toll-like receptors and cytokine receptors. TMA blunts TLR4 signalling in primary human hepatocytes and peripheral blood monocytic cells, and improves mouse survival after a lipopolysaccharide-induced septic shock. Consistent with this, genetic deletion and chemical inhibition of IRAK4 result in similar metabolic and immune improvements in HFD. In summary, TMA appears to be a key microbial compound inhibiting IRAK4 and mediating metabolic and immune effects with benefits upon HFD. Thereby we highlight the critical contribution of the microbial signalling metabolome in homeostatic regulation of host disease and the emerging role of the kinome in microbial-mammalian chemical crosstalk.

systems biology

ARG-miner: A web platform for crowdsourcing-based curation of antibiotic resistance genes

Curation of antibiotic resistance gene (ARG) databases is a labor-intensive process that requires expert knowledge to manually collect, correct, and/or annotate individual genes. Correspondingly, updates to existing databases tend to be infrequent, commonly requiring years for completion and often containing inconsistences. Further, because of limitations of manual curation, most existing ARG databases contain only a small proportion of known ARGs (~5k genes). A new approach is needed to achieve a truly comprehensive ARG database, while also maintaining a high level of accuracy. Here we propose a new web-based curation system, ARG-miner, which supports annotation of ARGs at multiple levels, including: gene name, antibiotic category, resistance mechanism, and evidence for mobility and occurrence in clinically-important bacterial strains. To overcome limitations of manual curation, we employ crowdsourcing as a novel strategy for expanding curation capacity towards achieving a truly comprehensive, up-to-date database. We develop and validate the approach by comparing performance of multiple cohorts of curators with varying levels of expertise, demonstrating that ARG-miner is more cost effective and less time-consuming relative to traditional expert curation. We further demonstrate the reliability of a trust validation filter for rejecting confounding input generated by spammers. Crowdsourcing was found to be as accurate as expert annotation, with an accuracy >90% for the annotation of a diverse test set of ARGs. ARG-miner provides a public API and database available at http://bench.cs.vt.edu/argminer.

bioinformatics

Learning common and specific patterns from data of multiple interrelated biological scenarios with matrix factorization

High-throughput biological technologies (e.g., ChIP-seq, RNA-seq and single-cell RNA-seq) rapidly accelerate the accumulation of genome-wide omics data in diverse interrelated biological scenarios (e.g., cells, tissues and conditions). Data dimension reduction and differential analysis are two common paradigms for exploring and analyzing such data. However, they are typically used in a separate or/and sequential manner. In this study, we propose a flexible non-negative matrix factorization framework CSMF to combine them into one paradigm to simultaneously reveal common and specific patterns from data generated under interrelated biological scenarios. We demonstrate the effectiveness of CSMF with four applications including pairwise ChIP-seq data describing the chromatin modification map on protein-DNA interactions between K562 and Huvec cell lines; pairwise RNA-seq data representing the expression profiles of two cancers (breast invasive carcinoma and uterine corpus endometrial carcinoma); RNA-seq data of three breast cancer subtypes; and single-cell sequencing data of human embryonic stem cells and differentiated cells at six time points. Extensive analysis yields novel insights into hidden combinatorial patterns embedded in these interrelated multi-modal data. Results demonstrate that CSMF is a powerful tool to uncover common and specific patterns with significant biological implications from data of interrelated biological scenarios.

bioinformatics

The proteome of the malaria plastid organelle, a key anti-parasitic target

Malaria parasites (Plasmodium spp.) and related apicomplexan pathogens contain a non-photosynthetic plastid called the apicoplast. Derived from an unusual secondary eukaryote-eukaryote endosymbiosis, the apicoplast is a fascinating organelle whose function and biogenesis rely on a complex amalgamation of bacterial and algal pathways. Because these pathways are distinct from the human host, the apicoplast is an excellent source of novel antimalarial targets. Despite its biomedical importance and evolutionary significance, the absence of a reliable apicoplast proteome has limited most studies to the handful of pathways identified by homology to bacteria or primary chloroplasts, precluding our ability to study the most novel apicoplast pathways. Here we combine proximity biotinylation-based proteomics (BioID) and a new machine learning algorithm to generate a high-confidence apicoplast proteome consisting of 346 proteins. Critically, the high accuracy of this proteome significantly outperforms previous prediction-based methods and extends beyond other BioID studies of unique parasite compartments. Half of identified proteins have unknown function, and 77% are predicted to be important for normal blood-stage growth. We validate the apicoplast localization of a subset of novel proteins and show that an ATP-binding cassette protein ABCF1 is essential for blood-stage survival and plays a previously unknown role in apicoplast biogenesis. These findings indicate critical organellar functions for newly discovered apicoplast proteins. The apicoplast proteome will be an important resource for elucidating unique pathways derived from secondary endosymbiosis and prioritizing antimalarial drug targets.

microbiology

Structure of the 30S ribosomal decoding complex at ambient temperature

The ribosome translates nucleotide sequences of messenger RNA to proteins through selection of cognate transfer RNA according to the genetic code. To date, structural studies of ribosomal decoding complexes yielding high-resolution data have predominantly relied on experiments performed at cryogenic temperatures. New lightsources like the X-ray free electron laser (XFEL) have enabled data collection from macromolecular crystals at ambient temperature. Here, we report an X-ray crystal structure of the Thermus thermophilus 30S ribosomal subunit decoding complex to 3.45 [A] resolution using data obtained at ambient temperature at the Linac Coherent Light Source (LCLS). We find that this ambient-temperature structure is largely consistent with existing cryogenic-temperature crystal structures, with key residues of the decoding complex exhibiting similar conformations, including adenosine residues 1492 and 1493. Minor variations were observed, namely an alternate conformation of cytosine 1397 near the mRNA channel and the A-site. Our serial crystallography experiment illustrates the amenability of ribosomal microcrystals to routine structural studies at ambient temperature, thus overcoming a long-standing experimental limitation.

biophysics

Reversing Glial Scar Back To Neural Tissue Through NeuroD1-Mediated Astrocyte-To-Neuron Conversion

Nerve injury often causes neuronal loss and glial proliferation, disrupting the delicate balance between neurons and glial cells in the brain. Recently, we have developed an innovative technology to convert internal reactive glial cells into functional neurons inside the mouse brain. Here, we further demonstrate that such glia-to-neuron conversion can rebalance neuron-glia ratio and reverse glial scar back to neural tissue. Specifically, using a severe stab injury model in the mouse cortex, we demonstrated that ectopic expression of NeuroD1 in reactive astrocytes significantly reduced glial reactivity and transformed toxic A1 astrocytes into less harmful astrocytes before neuronal conversion. Importantly, astrocytes were not depleted after neuronal conversion but rather repopulated due to its intrinsic proliferation capability. Remarkably, converting reactive astrocytes into neurons also significantly reduced microglia-mediated neuroinflammation. Moreover, accompanying regeneration of new neurons together with repopulation of new astrocytes, blood-brain-barrier was restored and synaptic density was rescued in the injury sites. Together, these results demonstrate that glial scar can be reversed back to neural tissue through rebalancing neuron:glia ratio after glia-to-neuron conversion.

neuroscience

DeepHINT: Understanding HIV-1 integration via deep learning with attention

MotivationHuman immunodeficiency virus type 1 (HIV-1) genome integration is closely related to clinical latency and viral rebound. In addition to human DNA sequences that directly interact with the integration machinery, the selection of HIV integration sites has also been shown to depend on the heterogeneous genomic context around a large region, which greatly hinders the prediction and mechanistic studies of HIV integration.\n\nResultsWe have developed an attention-based deep learning framework, named DeepHINT, to simultaneously provide accurate prediction of HIV integration sites and mechanistic explanations of the detected sites. Extensive tests on a high-density HIV integration site dataset showed that DeepHINT can outperform conventional modeling strategies by automatically learning the genomic context of HIV integration solely from primary DNA sequence information. Systematic analyses on diverse known factors of HIV integration further validated the biological relevance of the prediction result. More importantly, in-depth analyses of the attention values output by DeepHINT revealed intriguing mechanistic implications in the selection of HIV integration sites, including potential roles of several basic helix-loop-helix (bHLH) transcription factors and zinc-finger proteins. These results established DeepHINT as an effective and explainable deep learning framework for the prediction and mechanistic study of HIV integration.\n\nAvailabilityDeepHINT is available as an open-source software and can be downloaded from https://github.com/nonnerdling/DeepHINT\n\nContactlzhang20@mail.tsinghua.edu.cn and zengjy321@tsinghua.edu.cn

bioinformatics

From 1D sequence to 3D chromatin dynamics and cellular functions: a phase separation perspective

The high-order chromatin structure plays a non-negligible role in gene regulation. However, the mechanism for the formation of different chromatin structures in different cells and the sequence dependence of this process remain to be elucidated. As the nucleotide distributions in human and mouse genomes are highly uneven, we identified CGI forest and prairie genomic domains based on CGI density, which better segregates genomic elements along the genome than GC content. The genome is then divided into two sequentially, epigenetically, and transcriptionally distinct regions. These two types of megabase-sized domains spatially segregate, but to a different extent in different cell types. Overall, the forests and prairies gradually segregate from each other in development, differentiation, and senescence. The multi-scale forest-prairie spatial intermingling is cell-type specific and increases in differentiation, thus helps define the cell identity. We propose that the phase separation of the 1D mosaic sequence in space, serving as a potential driving force, together with cell type specific epigenetic marks and transcription factors, shapes the chromatin structure in different cell types and renders them distinct genomic properties. The mosaicity of the genome manifested in terms of alternative forests and prairies of a species could be related to its biological processes such as differentiation, aging and body temperature control.

biophysics

miRNAs play important roles in aroma weakening during the shelf life of ‘Nanguo’ pear after cold storage

Cold storage is commonly employed to delay senescence in Nanguo pears after harvest. However, this technique also causes fruit aroma weakening. MicroRNAs play important roles in plant development and in eliciting responses to abiotic environmental stressors. In this study, the miRNA transcript profile of the fruit at the first day (C0, LT0) move in and out of cold storage and the optimum tasting period (COTP, LTOTP) during shelf life at room temperature were analyzed, respectively. More than 300 known miRNAs were identified in Nanguo pears; 176 and 135 miRNAs were significantly differentially expressed on the C0 vs. LT0 and on the COTP vs. LTOTP, respectively. After prediction the target genes of these miRNAs, LOX2S, LOX1_5, HPL, and ADH1 were found differentially expressed, which were the key genes during aroma formation. The expression pattern of these target genes and the related miRNAs were identified by RT-PCR. Mdm-miR172a-h, mdm-miR159a/b/c, mdm-miR160a-e, mdm-miR395a-i, mdm/ppe-miR399a, mdm/ppe-miR535a/b, and mdm-miR7120a/b negatively regulated target gene expression. These results indicate that miRNAs play key roles in aroma weakening in refrigerated Nanguo pear and provide valuable information for studying the molecular mechanisms of miRNAs in the aroma weakening of fruits due to cold storage.

molecular biology

Identifying associations in dense connectomes using structured kernel principal component regression

A powerful and computationally efficient multivariate approach is proposed here, called structured kernel principal component regression (sKPCR), for the identification of associations in the voxel-level dense connectome. The method can identify voxel-phenotype associations based on the voxels whole-brain connectivity pattern, which is applicable to detect linear and non-linear signals for both volume-based and surface-based functional magnetic resonance imaging (fMRI) data. For each voxel, our approach first extracts signals from the spatially smoothed connectivities by structured kernel principal component analysis, and then tests the voxel-phenotype associations via a general linear model. The method derives its power by appropriately modelling the spatial structure of the data. Simulations based on dense connectome data have shown that our method can accurately control the false-positive rate, and it is more powerful than many state-of-the-art approaches, such as the connectivity-wise general linear model (GLM) approach, multivariate distance matrix regression (MDMR), adaptive sum of powered score (aSPU) test, and least-square kernel machine (LSKM). To demonstrate the utility of our approach in real data analysis, we apply these methods to identify voxel-wise difference between schizophrenic patients and healthy controls in two independent resting-state fMRI datasets. The findings of our approach have a better between-sites reproducibility, and a larger proportion of overlap with existing schizophrenia findings. Code for our approach can be downloaded from https://github.com/weikanggong/vBWAS.

neuroscience

Comparison of computational methods for imputing single-cell RNA-sequencing data

Single-cell RNA-sequencing (scRNA-seq) is a recent breakthrough technology, which paves the way for measuring RNA levels at single cell resolution to study precise biological functions. One of the main challenges when analyzing scRNA-seq data is the presence of zeros or dropout events, which may mislead downstream analyses. To compensate the dropout effect, several methods have been developed to impute gene expression since the first Bayesian-based method being proposed in 2016. However, these methods have shown very diverse characteristics in terms of model hypothesis and imputation performance. Thus, large-scale comparison and evaluation of these methods is urgently needed now. To this end, we compared eight imputation methods, evaluated their power in recovering original real data, and performed broad analyses to explore their effects on clustering cell types, detecting differentially expressed genes, and reconstructing lineage trajectories in the context of both simulated and real data. Simulated datasets and case studies highlight that there are no one method performs the best in all the situations. Some defects of these methods such as scalability, robustness and unavailability in some situations need to be addressed in future studies.

bioinformatics

Long-range infra-sound acoustic signaling inhuman in vivo

Acupuncture is widely deployed today, but its basic physiology with Qi and meridian is not understood. This letter postulates that Qi is an infrasound wave packet and meridian is the muscle waveguide. Using video cameras and signal processing in an IRB approved clinical experiment we performed a comparison between control and electro-activated statistical tests on the long range (50-80cm) unidirectional transmission of Qi (p = 0.025). In the reverse direction, there is no transmission even in a distance of less than 10 cm (p = 0.545). The rectification is a surprise but in full agreement with Huang Di Nei Jing.

physiology

Gene editing of the multi-copy H2A.B gene family by a single pair of TALENS

In view of the controversy related to the generation of off-target mutations by gene editing approaches, we tested the specificity of TALENs by disrupting a multi-copy gene family using only one pair of TALENS. We show here that TALENS do display a high level of specificity by simultaneously knocking out the function of the three genes that encode for H2A.B.3. This represents the first described knockout of this histone variant.

genomics