Search bioRxivSearch

Biology subjects

Jia, P.

Publications and source records attributed to Jia, P..

4 recordsLinked to original sources

A computational protocol to characterize elusive Candidate Phyla Radiation bacteria in oral environments using metagenomic data

Several studies have documented the diversity and potential pathogenic associations of organisms in the human oral cavity. Although much progress has been made in understanding the complex bacterial community inhabiting the human oral cavity, our understanding of some microorganisms is less resolved due to a variety of reasons. One such little-understood group is the candidate phyla radiation (CPR), which is a recently identified, but highly abundant group of ultrasmall bacteria with reduced genomes and unusual ribosomes. Here, we present a computational protocol for the detection of CPR organisms from metagenomic data. Our approach relies on a self-constructed dataset comprising published CPR genomic sequences as a filter to identify CPR sequences from metagenomic sequencing data. After assembly and functional prediction, the taxonomic affiliation of CPR contigs can be identified through phylogenetic analysis with publically available 16S rRNA gene and ribosomal proteins, in addition to sequence similarity analyses (e.g., average nucleotide identity calculations and contig mapping). Using this protocol, we reconstructed two draft genomes of organisms within the TM7 superphylum, that had genome sizes of 0.594 Mb and 0.678 Mb. Among the predicted functional genes of the constructed genomes, a high percentage were related to signal transduction, cell motility, and cell envelope biogenesis, which could contribute to cellular morphological changes in response to environmental cues.\n\nImportanceCandidate phyla radiation (CPR) bacterial group is a recently identified, but highly diverse and abundant group of ultrasmall bacteria exhibiting reduced genomes and limited metabolic capacities. A number of studies have reported their potential pathogenic associations in multiple mucosal diseases including periodontitis, halitosis, and inflammatory bowel disease. However, CPR organisms are difficult to cultivate and are difficult to detect with PCR-based methods due to divergent genetic sequences. Thus, our understanding of CPR has lagged behind that of other bacterial component. Here, we used metagenomic approaches to overcome these previous barriers to CPR identification, and established a computational protocol for detection of CPR organisms from metagenomic samples. The protocol describe herein holds great promise for better understanding the potential biological functioning of CPR. Moreover, the pipeline could be applied to other organisms that are difficult to cultivate.

bioinformatics

Differential DNA modification of an enhancer at the IGF2 locus affects dopamine synthesis in patients with major psychosis

Dopamine dysregulation is central to the pathogenesis of diseases with major psychosis, but its molecular origins are unclear. In an epigenome-wide investigation in neurons, individuals with schizophrenia and bipolar disorder showed reduced DNA modifications at an enhancer in IGF2, which disrupted the regulation of the dopamine synthesis enzyme tyrosine hydroxylase and striatal dopamine levels in transgenic mice. Epigenetic control of this enhancer may be an important molecular determinant of psychosis.

neuroscience

Gene2Vec: Distributed Representation of Genes Based on Co-Expression

Existing functional description of genes are categorical, discrete, and mostly through manual process. In this work, we explore the idea of gene embedding, distributed representation of genes, in the spirit of word embedding. From a pure data-driven fashion, we trained a 300 dimension vector representation of all human genes, using gene co-expression patterns in 984 data sets from the GEO databases. These vectors capture functional relatedness of genes in terms of recovering known pathways - the average inner product (similarity) of genes within a pathway is 1.68X greater than that of random genes. Using t-SNE, we produced a gene co-expression map that shows local concentrations of tissue specific genes. We also illustrated the usefulness of the embedded gene vectors, laden with rich information on gene co-expression patterns, in tasks such as gene-gene interaction prediction. Overall, we believe that this distributed representation of genes may be useful for more bioinformatics applications.

bioinformatics

scRNASeqDB: a database for gene expression profiling in human single cell by RNA-seq

Summary: Single-cell RNA sequencing (scRNA-Seq) is quickly becoming a powerful tool for high-throughput transcriptomic analysis of cell states and dynamics. Both the number and quality of scRNA-Seq datasets have dramatically increased recently. So far, there is no database that comprehensively collects and curates scRNA-Seq data in humans. Here, we present scRNASeqDB, a database that includes almost all the currently available human single cell transcriptome datasets (n= 36) covering 71 human cell lines or types and 8910 samples. Our online web interface allows user to query and visualize expression profiles of the gene(s) of interest, search for genes that are expressed in different cell types or groups, or retrieve differentially expressed genes between cell types or groups. The scRNASeqDB is a valuable resource for single cell transcriptional studies.\n\nAvailability: The database is available at https://bioinfo.uth.edu/scrnaseqdb/.\n\nContact: zhongming.zhao@uth.tmc.edu

bioinformatics