Search bioRxivSearch

Biology subjects

Chen, X.

Publications and source records attributed to Chen, X..

At least 73 records · Page 4Linked to original sources

Joint single-cell DNA accessibility and protein epitope profiling reveals environmental regulation of epigenomic heterogeneity

Here we introduce Protein-indexed Assay of Transposase Accessible Chromatin with sequencing (Pi-ATAC) that combines single-cell chromatin and proteomic profiling. In conjunction with DNA transposition, the levels of multiple cell surface or intracellular protein epitopes are recorded by index flow cytometry and positions in arrayed microwells, and then subject to molecular barcoding for subsequent pooled analysis. Pi-ATAC simultaneously identifies the epigenomic and proteomic heterogeneity in individual cells. Pi-ATAC reveals a casual link between transcription factor abundance and DNA motif access, and deconvolute cell types and states in the tumor microenvironment in vivo. We identify a dominant role for hypoxia, marked by HIF1 protein, in the tumor microvenvironment for shaping the regulome in a subset of epithelial tumor cells.

genomics

Epitope-based vaccine design yields fusion peptide-directed antibodies that neutralize diverse strains of HIV-1

A central goal of HIV-1-vaccine research is the elicitation of antibodies capable of neutralizing diverse primary isolates of HIV-1. Here we show that focusing the immune response to exposed N-terminal residues of the fusion peptide, a critical component of the viral entry machinery and the epitope of antibodies elicited by HIV-1 infection, through immunization with fusion peptide-coupled carriers and prefusion-stabilized envelope trimers, induces cross-clade neutralizing responses. In mice, these immunogens elicited monoclonal antibodies capable of neutralizing up to 31% of a cross-clade panel of 208 HIV-1 strains. Crystal and cryo-electron microscopy structures of these antibodies revealed fusion peptide-conformational diversity as a molecular explanation for the cross-clade neutralization. Immunization of guinea pigs and rhesus macaques induced similarly broad fusion peptide-directed neutralizing responses suggesting translatability. The N terminus of the HIV-1-fusion peptide is thus a promising target of vaccine efforts aimed at eliciting broadly neutralizing antibodies.

immunology

LCA robustly reveals subtle diversity in large-scale single-cell RNA-seq data

Single-cell RNA sequencing has emerged as a powerful tool for characterizing the cell-to-cell variation and dynamics. We present Latent Cellular Analysis (LCA), a machine learning-based analytical pipeline that features a dual-space model search with inference of latent cellular states, control of technical variations, cosine similarity measurement, and spectral clustering. LCA has proved to be robust, accurate, scalable, and powerful in revealing subtle diversity in cell populations.

bioinformatics

DNA local structure decreases mutation rates

BackgroundMutation rates vary across the genome. Whereas many trans factors that influence mutation rates have been identified, as have specific sequence motifs at the 1-7 bp scale, cis elements remain poorly characterized. The lack of understanding why different sequences have different mutation rates hampers our ability to identify positive selection in evolution and to identify driver mutations in tumorigenesis.\n\nResultsHere we show, using a combination of synthetic genes and sequencing of thousands of isolated yeast colonies, that intrinsic DNA curvature is the major cis determinant of mutation rate. Mutation rate negatively correlates with DNA curvature within genes, and a 10% decrease in curvature results in a 70% increase in mutation rate. Consistently, both yeast cells and human tumors accumulate mutations in regions with small curvature. We further show that this effect is due to differences in the intrinsic mutation rate, likely due to differences in mutagen sensitivity, and not due to differences in the local activity of DNA repair.\n\nConclusionsOur study establishes a framework in understanding the cis properties of DNA sequence in modulating the local mutation rate and identifies a novel causal source of non-uniform mutation rates across the genome.

genetics

Spatial organization of projection neurons in the mouse auditory cortex identified by in situ barcode sequencing

Understanding neural circuits requires deciphering interactions among myriad cell types defined by spatial organization, connectivity, gene expression, and other properties. Resolving these cell types requires both single neuron resolution and high throughput, a challenging combination with conventional methods. Here we introduce BARseq, a multiplexed method based on RNA barcoding for mapping projections of thousands of spatially resolved neurons in a single brain, and relating those projections to other properties such as gene or Cre expression. Mapping the projections to 11 areas of 3579 neurons in mouse auditory cortex using BARseq confirmed the laminar organization of the three top classes (IT, PT-like and CT) of projection neurons. In depth analysis uncovered a novel projection type restricted almost exclusively to transcriptionally-defined subtypes of IT neurons. By bridging anatomical and transcriptomic approaches at cellular resolution with high throughput, BARseq can potentially uncover the organizing principles underlying the structure and formation of neural circuits.

neuroscience

Astroplastic: A start-to-finish process for polyhydroxybutyrate production from solid human waste using genetically engineered bacteria to address the challenges for future manned Mars missions

Space exploration has long been a source of inspiration, challenging scientists and engineers to find innovative solutions to various problems. One of the current focuses in space exploration is to send humans to Mars. However, the challenge of transporting materials to Mars and the need for waste management processes are two major obstacles for these long-duration missions.\n\nTo address these two challenges a process called Astroplastic was developed that produces polyhydroxybutyrate (PHB) from solid human waste, which can be used to 3D print useful items for astronauts. PHB granules are naturally produced by bacteria such as Ralstonia eutropha and Pseudomonas aeruginosa for carbon and energy storage. The phaJ, phaC, and phaCBA genes were cloned from these native PHB-producing bacteria into Escherichia coli. These genes code for enzymes that aid in PHB production by converting products of glycolysis and {beta}-oxidation pathways, such as acetyl-CoA and enoyl-CoA, into PHB. To ensure a continuous PHB production system and to eliminate the need for cell lysis to extract PHB, recombinant E. coli was engineered to use the genes in its natural type I secretion system to secrete PHB. The C-terminal of the HlyA secretion tag was fused to phasin (PhaP), a protein originally from R. eutropha. Phasin-HlyA electrostatically binds PHB granules and transports them outside of the cell.\n\nIn addition to genetically engineering bacteria, a concept for start-to-finish PHB production process was designed. Integrating expert feedback and experimental results, conditions for each step of the process including the collection and storage of waste, volatile fatty acid (VFA) fermentation, VFA extraction, PHB fermentation, and PHB extraction were optimized. The optimized system will provide a sustainable and continuous PHB production system, which will address the problems of transportation costs and waste management for future space missions.\n\nFinancial DisclosureMindfuel Science Alberta Foundation Genome Alberta GenScript Polyferm Canada GeekStarter Alberta Integrated DNA Technologies University of Calgary University of Calgary Cumming School of Medicine University of Calgary Bachelor of Sciences University of Calgary Schulich School of Engineering University of Calgary OBrien Centre for the Bachelor of Health Sciences City of Calgary Alberta Innovates The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.\n\nCompeting InterestsThe authors have declared that no competing interests exist.\n\nEthics StatementN/A\n\nData AvailabilityAll data are freely available without restriction.

synthetic biology

Widespread and polymorphous noncoding amino acid residues in human sperm proteome

Proteins are usually deciphered by translation of the coding genome; however, their amino acid residues are seldom determined directly across the proteome. Herein, we describe a systematic workflow for identifying all possible protein residues that differ from the coding genome, termed noncoded amino acids (ncAAs). By measuring the mass differences between the coding amino acids and the actual protein residues in human spermatozoa, over a million nonzero delta masses were detected, fallen into 424 high-quality Gaussian clusters and 571 high-confidence ncAAs spanning 29,053 protein sites. Most ncAAs are novel with unresolved side-chains and discriminative between healthy individuals and patients with oligoasthenospermia. For validation, 40 out of 98 ncAAs that matched with amino acid substitutions were confirmed by exon sequencing. This workflow revealed the widespread existence of previously unreported ncAAs in the sperm proteome, which represents a new dimension on the understanding of amino acid polymorphisms at the proteomic level.\n\nHighlightsO_LI571 ncAAs spanning 108,000 protein sites were identified in human sperm proteome.\nC_LIO_LIMost ncAAs are novel with unresolved sidechains and found at unreported protein sites.\nC_LIO_LIExon sequencing confirmed 40 of 98 ncAAs that matched with amino acid substitutions.\nC_LIO_LIMany ncAAs are linked with disease and have potential for diagnosis and targeting.\nC_LI\n\neTOC BlurbWe describe a systematic identification of all possible protein residues that were not encoded by their genomic sequences. A total of 571 high-confidence most novel noncoded amino acids were identified in human sperm proteome, corresponding to over 108,000 ncAA-containing protein sites. For validation, 40 out of 98 ncAAs that matched to amino acid substitutions were confirmed by exon sequencing. These ncAAs are discriminative between individuals and expand our understanding of amino acid polymorphisms in human proteomes and diseases.

molecular biology

Dosage sensitivity of X-linked genes in human embryonic single cells

Fifty years ago, Susumu Ohno proposed that the expression levels of X-linked genes have doubled as dosage compensation for autosomal genes due to degeneration of Y-linked homologs during evolution of mammalian sex chromosomes. Recent studies have nevertheless shown that the X to autosome expression ratio equals ~1 in haploid human parthenogenetic embryonic stem (pES) cells and ~0.5 in diploid pES cells, thus refuting Ohnos hypothesis. Here, by reanalyzing a RNA-seq-based single-cell transcriptome dataset of human embryos (Petropoulos, et al. 2016), we found that from the 8-cell stage until the time-point just prior to implantation, the expression levels of X-linked genes are not two-fold upregulated in male cells and gradually decrease from two-fold in female cells. This observation suggests that the expression levels of X-linked genes are imbalanced, with autosomal genes starting from the early 8-cell stage, and that the dosage conversion is fast, such that the X:AA expression ratio reaches ~0.5 in no more than a week. Additional analyses of gene expression noise further suggest that the dosage sensitivity of X-linked genes is weaker than that of autosomal genes in differentiated female cells, which contradicts a key assumption of Ohnos hypothesis. Moreover, the dosage-sensitive housekeeping genes are preferentially located on autosomes, implying selection against X-linkage for dosage-sensitive genes. Our results collectively suggest an alternative to Ohnos hypothesis that X-linked genes are less likely to be dosage sensitive than autosomal genes.

genomics

Brainwide organization of neuronal activity and convergent sensorimotor transformations in larval zebrafish

Simultaneous recordings of large populations of neurons in behaving animals allow detailed observation of high-dimensional, complex brain activity. However, experimental design and analysis approaches have not sufficiently evolved to fully realize the potential of these methods. We recorded whole-brain neuronal activity for larval zebrafish presented with a battery of visual stimuli while recording fictive motor output. These data were used to develop analysis methods including regression techniques that leverage trial-to-trial variations and unsupervised clustering techniques that organize neurons into functional groups. We used these methods to obtain brain-wide maps of concerted activity, which revealed both known and heretofore uncharacterized brain nuclei. We also identified neurons tuned to each stimulus type and motor output, and revealed nuclei in the anterior hindbrain that respond to multiple stimuli that elicit the same behavior. However, these convergent sensorimotor representations were only weakly correlated to instantaneous motor behavior, suggesting that they inform, but do not directly generate, behavioral output. These findings motivate a novel model of sensorimotor transformation spanning distinct behavioral contexts, within which these hindbrain convergence neurons likely constitute a key step.

neuroscience

Apparent oxygen half saturation constant for nitrifiers: genus specific, inherent physiological property, or artefact of colony morphology?

We report that a single Nitrospira sublineage I OTU performs nitrite oxidation in several full-scale domestic wastewater treatment plants (WWTPs) in the tropics (29-31 {degrees}C). Contrary to the prevailing theory for the relationship between nitrite oxidizing bacteria (NOB) and ammonia oxidizing bacteria (AOB), members of the Nitrospira sublineage I OTU had an apparent half saturation coefficient, Ks(app) lower than that of the full-scale domestic activated sludge cohabitant AOB (0.09 {+/-} 0.02 g O2 m-3 versus 0.3 {+/-} 0.03 g O2 m-3). Paradoxically, NOB may thus thrive under conditions of low oxygen supply. Low dissolved oxygen (DO) conditions could enrich for and high aeration inhibit the NOB in a long-term lab-scale reactor. The relative abundance of Nitrospira gradually decreased with increasing DO until it was washed out. Nitritation was sustained even after the DO was lowered subsequently. Based on 3D-fluorescence in situ hybridization (FISH) image analysis, the morphologies of AOB and NOB microcolonies responded to DO levels in accordance with their apparent oxygen half saturation constant Ks(app). When exposed to the same oxygenation level, NOB formed densely packed spherical clusters with a low surface area-to-volume ratio compared to the Nitrosomonas-like AOB clusters, which maintained a porous and non-spherical morphology. Microcolony morphology is thus a way for AOB and NOB to regulate oxygen exposure and sustain the mutualistic interaction. However, short-term high DO exposure can select for AOB and against NOB in full-scale domestic WWTPs and such population dynamics depend on which specific AOB and NOB species predominate under given environmental conditions.

systems biology

WDR45 contributes to neurodegeneration through regulation of ER homeostasis and neuronal death

Mutations in the autophagy gene WDR45 cause {beta}-propeller protein-associated neurodegeneration (BPAN); however the molecular and cellular mechanism of the disease process is largely unknown. Here we generated constitutive Wdr45 knockout (KO) mice that displayed cognitive impairments, abnormal synaptic transmission and lesions in hippocampus and basal ganglia. Immunohistochemistry analysis shows loss of neurons in prefrontal cortex and basal ganglion in aged mice, and increased apoptosis in these regions, recapitulating a hallmark of neurodegeneration. Quantitative proteomic analysis shows accumulation of endoplasmic reticulum (ER) proteins in KO mouse. Furthermore, we show that a defect in autophagy results in impaired ER turnover and ER stress. The unfolded protein response (UPR) is elevated through IRE1 and possibly other kinase signaling pathways, and eventually leads to neuronal apoptosis. Suppression of ER stress, or activation of autophagy through inhibition of mTOR pathway rescues neuronal death. Thus, our study not only provides mechanistic insights for BPAN, but also suggests that a defect in macroautophagy machinery leads to impairment in selective organelle autophagy.

cell biology

Physiological Significance of R-fMRI Indices in Detecting Structural Brain Lesions

Resting-state functional MRI (R-fMRI) research has recently entered the era of \"big data\", however, few studies have provided a rigorous validation of the physiological underpinnings of R-fMRI indices. Although studies have reported that various neuropsychiatric disorders exhibit abnormalities in R-fMRI measures, these \"biomarkers\" have not been validated in differentiating structural lesions (brain tumors) as a concept proof. We enrolled 60 patients with intracranial tumors located in the unilateral cranial cavity and 60 matched normal controls to test whether R-fMRI indices can differentiate tumors, which represents a prerequisite for adapting such indices as biomarkers for neuropsychiatric disorders. Common R-fMRI indices of tumors and their counterpart control regions, which were defined as the contralateral normal areas (for amplitude of low frequency fluctuations (ALFF), fractional ALFF (fALFF), regional homogeneity (ReHo) and degree centrality (DC)) and ipsilateral regions surrounding the tumors (for voxel-mirrored homotopic connectivity (VMHC)), were comprehensively assessed. According to paired t-tests with a Bonferroni correction, only ALFF (both with and without Z-standardization) and VMHC (Fishers r-to-z transformed) could successfully differentiate substantial tumors from their counterpart normal regions in patients. And DC was not able to differentiate tumor from normal unless employed Z-standardization. To validate the lower power in the between-subject design than in the within-subject design, each metric was calculated in a matched control group, and two-sample t-tests were used to compare the patient tumors and the normal controls at the same area. Only ALFF (and that with Z-standardization) along with VMHC succeeded in differentiating significant differences between tumors and the sham tumors areas of normal controls. This study tested the premise of R-fMRI biomarkers for differentiating lesions, and brings a new understanding to physical significance of the Z-standardization.

neuroscience

Differential expression of an alternative splice variant of IL-12Rβ1 impacts early dissemination in the mouse and associates with disease outcome in both mouse and humans exposed to tuberculosis

Experimental mouse models of TB suggest that early events in the lung impact immunity. Early events in the human lung in response to TB are difficult to probe and their impact on disease outcome is unknown. We have shown in mouse that a secreted alternatively-spliced variant of IL-12R{beta}1, lacking the transmembrane domain and termed {Delta}TM-IL-12R{beta}1, promotes dendritic cell migration to the draining lymph node, augments T cell activation and limits dissemination of M. tuberculosis (Mtb). We show here that CBA/J and C3H/HeJ mice (both highly susceptible to Mtb) express higher levels of {Delta}TM-IL-12R{beta}1 than resistant C57BL6 mice and limit early dissemination of Mtb from the lungs. Both CD11c+ cells and T cells express {Delta}TM-IL-12R{beta}1 in humans, and mice unable to make {Delta}TM-IL-12R{beta}1 in either CD4 or CD11c expressing cells permit early dissemination from the lung. Analysis of publically available blood transcriptomes indicates that pulmonary TB is associated with high {Delta}TM-IL-12R{beta}1 expression and that of all IL-12 related signals, the {Delta}TM-IL-12R{beta}1 signal best predicts active disease. {Delta}TM-IL-12R{beta}1 expression reflects the heterogeneity of latent TB infection and has the capacity to discriminate between latent and active disease. In a new Chinese TB patient cohort, {Delta}TM-IL-12R{beta}1 effectively differentiates TB from latent TB, healthy controls and pneumonia patients. Finally, {Delta}TM-IL-12R{beta}1 expression drops in drug-treated individuals in the UK and China where infection pressure is low. We propose that {Delta}TM-IL-12R{beta}1 regulates early dissemination from the lung and that it has diagnostic potential and provides mechanistic insights into human TB.

immunology

Metabolomics and proteomics analyses of grain yield reduction in rice under abrupt drought-flood alternation

HighlightAbrupt drought-flood alteration is a frequent meteorological disaster that occurs during summer in southern China and the Yangtze river basin, which often causes a large area reduction of rice yield. We previously reported abrupt drought-flood alteration effects on yield and its components, physiological characteristics, matter accumulation and translocation, rice quality of rice. However, the molecular mechanism of rice yield reduction caused by abrupt drought-flood alternation has not been reported.\n\nIn this study, four treatments were provided, no drought and no floods (control), drought without floods (duration of drought 10 d), no drought with floods (duration of floods 8 d), and abrupt drought-flood alteration (duration of drought 10 d and floods 8 d). The quantitative analysis of spike metabolites was proceeded by LC-MS (liquid chromatograph-mass spectrometry) firstly. Then the Heat-map, PCA, PLS-DA, OPLS-DA and response ranking test of OPLS-DA model methods were used to analysis the function of differential metabolites (DMs) during the rice panicle differentiation stage under abrupt drought-flood alteration. In addition, relative quantitative analysis of spike total proteins under the treatment was conducted iTRAQ (isobaric tags for relative and absolute quantification) and LC-MS. In this study, 5708 proteins were identified and 4803 proteins were quantified. The identification and analysis of DEPs function suggested that abrupt drought-flood alteration treatment can promote carbohydrate metabolic, stress response, oxidation-reduction, defense response, and energy reserve metabolic process, etc, during panicle differentiation stage. In this study relative quantitative proteomics, metabolomics and physiology data (soluble protein content, superoxide dismutase activity, hydrogen peroxidase activity, peroxidase activity, malondialdehyde content, free proline content, soluble sugar content and net photosynthetic rate) analysis were applied to explicit the response mechanism of rice panicle differentiation stage under abrupt drought-flood alteration and provides a theoretical basis for the disaster prevention and mitigation.\n\nAbstractAbrupt drought-flood alternation is a meteorological disaster that frequently occurs during summer in southern China and the Yangtze river basin, often causing a significant loss of rice production. In this study, a quantitative analysis of spike metabolites was conducted via liquid chromatograph-mass spectrometry (LC-MS), and Heat-map, PCA, PLS-DA, OPLS-DA, and a response ranking test of OPLS-DA model methods were used to analyze functions of differential metabolites (DMs) during the rice panicle differentiation stage under abrupt drought-flood alternation. The results showed that 102 DMs were identified from the rice spike between T1 (abrupt drought-flood alternation) and CK0 (control) treatment, 104 DMs were identified between T1 and CK1 (drought) treatment and 116 DMs were identified between T1 and CK2 (flood) treatment. In addition, a relative quantitative analysis of spike total proteins was conducted using isobaric tags for relative and absolute quantification (iTRAQ) and LC-MS. The identification and analysis of DEPs functions indicates that abrupt drought-flood alternation treatment can promote carbohydrate metabolic, stress response, oxidation-reduction, defense response, and energy reserve metabolic process during the panicle differentiation stage. In this study, relative quantitative metabolomics and proteomics analyses were applied to explore the response mechanism of rice panicle differentiation in response to abrupt drought-flood alternation.\n\nAbbreviations

plant biology

VirTect: a computational method for detecting virus species from RNA-Seq and its application in head and neck squamous cell carcinoma

Next generation sequencing (NGS) provides an opportunity to detect viral species from RNA-seq data on human tissues, but existing computational approaches do not perform optimally on clinical samples. We developed a bioinformatics method called VirTect for detecting viruses in neoplastic human tissues using RNA-seq data. Here, we used VirTect to analyze RNA-seq data from 363 HNSCC (head and neck squamous cell carcinoma) patients and identified 22 HPV-induced HNSCCs. These predictions were validated by manual review of pathology reports on histopathologic specimens. Compared to two existing prediction methods, VirusFinder and VirusSeq, VirTect demonstrated superior performance with many fewer false positives and false negatives. The majority of HPV carcinogenesis studies thus far have been performed on cervical cancer and generalized to HNSCC. Our results suggest that HPV-induced HNSCC involves unique mechanisms of carcinogenesis, so understanding these molecular mechanisms will have a significant impact on therapeutic approaches and outcomes. In summary, VirTect can be an effective solution for the detection of viruses with NGS data, and can facilitate the clinicopathologic characterization of various types of cancers with broad applications for oncology.\n\nSignificance StatementWe developed a new bioinformatics tool, and reported the new inside of HPV carcinogenesis mechanism in HPV-induced head and neck squamous cell carcinoma (HNSCC). This novel bioin-formatics tool and the new knowledge of HPV-induced HNSCC will facilitate the development of target therapies for treating HNSCC.

bioinformatics

Programmable single and multiplex base-editing in Bombyx mori using RNA-guided cytidine deaminases

Standard genome editing tools (ZFN, TALEN and CRISPR/Cas9) edited genome depending on DNA double strand breaks (DSBs). A series of new CRISPR tools that convert cytidine to thymine (C to T) without the requirement for DNA double-strand breaks were developed recently, which have changed this status and have been quickly applied in a variety of organisms. Here, we demonstrate that CRISPR/Cas9-dependent base editor (BE3) converts C to T with a high frequency in the invertebrate Bombyx mori silkworm. Using BE3 as a knock-out tool, we inactivated exogenous and endogenous genes through base-editing-induced nonsense mutations with an efficiency of up to 66.2%. Furthermore, genome-scale analysis showed that 96.5% of B. mori genes have one or more targetable sites being knocked out by BE3 with a median of 11 sites per gene. The editing window of BE3 reached up to 13 bases (from C1 to C13 in the range of gRNA) in B. mori. Notably, up to 14 bases were substituted simultaneously in a single DNA molecule, with a low indel frequency of 0.6%, when 32 gRNAs were co-transfected. Collectively, our data show for the first time that RNA-guided cytidine deaminases are capable of programmable single and multiplex base-editing in an invertebrate model.

genetics

GeneQC: A quality control tool for gene expression estimation based on RNA-sequencing reads mapping

MotivationOne of the main benefits of using modern RNA-sequencing (RNA-Seq) technology is the more accurate gene expression estimations compared with previous generations of expression data, such as the microarray. However, numerous issues can result in the possibility that an RNA-Seq read can be mapped to multiple locations on the reference genome with the same alignment scores, which occurs in plant, animal, and metagenome samples. Such a read is so-called a multiple-mapping read (MMR). The impact of these MMRs is reflected in gene expression estimation and all downstream analyses, including differential gene expression, functional enrichment, etc. Current analysis pipelines lack the tools to effectively test the reliability of gene expression estimations, thus are incapable of ensuring the validity of all downstream analyses.\n\nResultsOur investigation into 95 RNA-Seq datasets from seven species (totaling 1,951GB) indicates an average of roughly 22% of all reads are MMRs for plant and animal species. Here we present a tool called GeneQC (Gene expression Quality Control), which can accurately estimate the reliability of each genes expression level. The underlying algorithm is designed based on extracted genomic and transcriptomic features, which are then combined using elastic-net regularization and mixture model fitting to provide a clearer picture of mapping uncertainty for each gene. GeneQC allows researchers to determine reliable expression estimations and conduct further analysis on the gene expression that is of sufficient quality. This tool also enables researchers to investigate continued re-alignment methods to determine more accurate gene expression estimates for those with low reliability.\n\nAvailabilityGeneQC is freely available at http://bmbl.sdstate.edu/GeneQC/home.html.\n\nContactqin.ma@sdstate.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Computational elucidation of regulatory network responding to acid stress in Lactococcus lactis MG1363

Acid stress caused by lactate increment can lead to the growth inhibition of bacteria and yes has not been fully defined. Regulons, serve as co-regulated gene groups contribute to the transcriptional regulation of microbe genome, have the potential in understanding the underlying regulatory mechanism. Lactococcus lactis is one of the most important Gram-positive lactic acid-producing bacteria, widely used in food industry and has been proved to have advantages in oral delivery of drug and vaccine. In this study, we designed a novel computational pipeline, RECTA, for regulon prediction. The pipeline carried out differentially expressed gene prediction, gene co-expression analysis, cis-regulatory motif finding, and comparative genomic study to predict and validate regulons related to acid stress response in Lactococcus lactis MG1363. A total of 51 regulons were identified, and 14 of them have computational verified significance. Among these 14 regulons, five of them were computationally predicted to be connected with acid stress response with (i) known transcriptional factors in MEME suite database successfully mapped in Lactococcus lactis MG1363; and (ii) differentially expressed genes between pH values of 6.5 (control) and 5.1 (treatment). Validated by 36 literature confirmed acid stress response related proteins and genes, 33 genes in Lactococcus lactis MG1363 were found having orthologous genes using BLAST, associated to six regulons. An acid response related regulatory network was constructed, involving two trans-membrane proteins, eight regulons (llrA, llrC, hllA, ccpA, NHP6A, rcfB, regulons #8 and #39), nine functional modules, and 33 genes with orthologous genes known to be associated to acid stress. Our RECTA pipeline provides an effective way to construct a reliable gene regulatory network based on regulon elucidation. The predicted resistance pathways could serve as promising candidates for better acid tolerance engineering in Lactococcus lactis. It has a strong application power and can be effectively applied to other bacterial genomes, where the elucidation of the transcriptional regulation network is needed.

systems biology