Search bioRxivSearch

Biology subjects

Yu, P.

Publications and source records attributed to Yu, P..

8 recordsLinked to original sources

A data mining paradigm for identifying key factors in biological processes using gene expression data

A large volume of biological data is being generated for studying mechanisms of various biological processes. These precious data enable large-scale computational analyses to gain biological insights. However, it remains a challenge to mine the data efficiently for knowledge discovery. The heterogeneity of these data makes it difficult to consistently integrate them, slowing down the process of biological discovery. We introduce a data processing paradigm to identify key factors in biological processes via systematic collection of gene expression datasets, primary analysis of data, and evaluation of consistent signals. To demonstrate its effectiveness, our paradigm was applied to epidermal development and identified many genes that play a potential role in this process. Besides the known epidermal development genes, a substantial proportion of the identified genes are still not supported by gain- or loss-of-function studies, yielding many novel genes for future studies. Among them, we selected a top gene for loss-of-function experimental validation and confirmed its function in epidermal differentiation, proving the ability of this paradigm to identify new factors in biological processes. In addition, this paradigm revealed many key genes in cold-induced thermogenesis using data from cold-challenged tissues, demonstrating its generalizability. This paradigm can lead to fruitful results for studying molecular mechanisms in an era of explosive accumulation of publicly available biological data.

bioinformatics

RBPMetaDB: A comprehensive annotation of mouse RNA-Seq datasets with perturbations of RNA-binding proteins

RNA-binding proteins may play a critical role in gene regulation in various diseases or biological processes by controlling post-transcriptional events such as polyadenylation, splicing, and mRNA stabilization via binding activities to RNA molecules. Due to the importance of RNA-binding proteins in gene regulation, a great number of studies have been conducted, resulting in a large amount of RNA-Seq datasets. However, these datasets usually do not have structured organization of metadata, which limits their potentially wide use. To bridge this gap, the metadata of a comprehensive set of publicly available mouse RNA-Seq datasets with perturbed RNA-binding proteins were collected and integrated into a database called RBPMetaDB. This database contains 278 mouse RNA-Seq datasets for a comprehensive list of 163 RNA-binding proteins. These RNA-binding proteins account for only [~]10% of all known RNA-binding proteins annotated in Gene Ontology, indicating that most are still unexplored using high-throughput sequencing. This negative information provides a great pool of candidate RNA-binding proteins for biologists to conduct future experimental studies. In addition, we found that DNA-binding activities are significantly enriched among RNA-binding proteins in RBPMetaDB, suggesting that prior studies of these DNA- and RNA-binding factors focus more on DNA-binding activities instead of RNA-binding activities. This result reveals the opportunity to efficiently reuse these data for investigation of the roles of their RNA-binding activities. A web application has also been implemented to enable easy access and wide use of RBPMetaDB. It is expected that RBPMetaDB will be a great resource for improving understanding of the biological roles of RNA-binding proteins.\n\nDatabase URL: http://rbpmetadb.yubiolab.org

bioinformatics

Genome-wide transcriptome analysis identifies alternative splicing regulatory network and key splicing factors in mouse and human psoriasis

Psoriasis is a chronic inflammatory disease that affects the skin, nails, and joints. For understanding the mechanism of psoriasis, though, alternative splicing analysis has received relatively little attention in the field. Here, we developed and applied several computational analysis methods to study psoriasis. Using psoriasis mouse and human datasets, our differential alternative splicing analyses detected hundreds of differential alternative splicing changes. Our analysis of conservation revealed many exon-skipping events conserved between mice and humans. In addition, our splicing signature comparison analysis using the psoriasis datasets and our curated splicing factor perturbation RNA-Seq database, SFMetaDB, identified nine candidate splicing factors that may be important in regulating splicing in the psoriasis mouse model dataset. Three of the nine splicing factors were confirmed upon analyzing the human data. Our computational methods have generated predictions for the potential role of splicing in psoriasis. Future experiments on the novel candidates predicted by our computational analysis are expected to provide a better understanding of the molecular mechanism of psoriasis and to pave the way for new therapeutic treatments.

bioinformatics

GEOMetaCuration: A web-based application for accurate manual curation of Gene Expression Omnibus metadata

Metadata curation has become increasingly important for biological discovery and biomedical research because a large amount of heterogeneous biological data is currently freely available. To facilitate efficient metadata curation, we developed an easy-to-use web-based curation application, GEOMetaCuration, for curating the metadata of Gene Expression Omnibus datasets. It can eliminate mechanical operations that consume precious curation time and can help coordinate curation efforts among multiple curators. It improves the curation process by introducing various features that are critical to metadata curation, such as a back-end curation management system and a curator-friendly front-end. The application is based on a commonly used web development framework of Python/Django and is open-sourced under the GNU General Public License V3. GEOMetaCuration is expected to benefit the biocuration community and to contribute to computational generation of biological insights using large-scale biological data. An example use case can be found at the demo website: http://geometacuration.yubiolab.org. Source code URL: https://bitbucket.com/yubiolab/GEOMetaCuration

bioinformatics

CITGeneDB: A comprehensive database of human and mouse genes enhancing or suppressing cold-induced thermogenesis validated by perturbation experiments in mice

Cold-induced thermogenesis increases energy expenditure and can reduce body weight in mammals, so the genes involved in it are thought to be potential therapeutic targets for treating obesity and diabetes. In the quest for more effective therapies, a great deal of research has been conducted to elucidate the regulatory mechanism of cold-induced thermogenesis. Over the last decade, a large number of genes that can enhance or suppress cold-induced thermogenesis have been discovered, but a comprehensive list of these genes is lacking. To fill this gap, we examined all of the annotated human and mouse genes and curated those demonstrated to enhance or suppress cold-induced thermogenesis by in vivo or ex vivo experiments in mice. The results of this highly accurate and comprehensive annotation are hosted on a database called CITGeneDB, which includes a searchable web interface to facilitate broad public use. The database will be updated as new genes are found to enhance or suppress cold-induced thermogenesis. It is expected that CITGeneDB will be a valuable resource in future explorations of the molecular mechanism of cold-induced thermogenesis, helping pave the way for new obesity and diabetes treatments. Database URL: http://citgenedb.yubiolab.org

bioinformatics

Root type and soil phosphate determine the taxonomic landscape of colonizing fungi and the transcriptome of field-grown maize roots

Key findingOur data illustrates for the first time that root type identity and phosphate availability determine the community composition of colonizing fungi and shape the transcriptomic response of the maize root system.\n\nSummaryO_LIPlant root systems consist of different root types colonized by a myriad of soil microorganisms including fungi, which influence plant health and performance. The distinct functional and metabolic characteristics of these root types may influence root type inhabiting fungal communities.\nC_LIO_LIWe performed internal transcribed spacer (ITS) DNA profiling to determine the composition of fungal communities in field-grown axial and lateral roots of maize (Zea mays L.) and in response to two different soil phosphate (P) regimes. In parallel, these root types were subjected to transcriptome profiling by RNA-Seq.\nC_LIO_LIWe demonstrated that fungal communities were influenced by soil P levels in a root type-specific manner. Moreover, maize transcriptome sequencing revealed root type-specific shifts in cell wall metabolism and defense gene expression in response to high phosphate. Furthermore, lateral roots specifically accumulated defense related transcripts at high P levels. This observation was correlated with a shift in fungal community composition including a reduction of colonization by arbuscular mycorrhiza fungi as observed in ITS sequence data and microscopic evaluation of root colonization.\nC_LIO_LIOur findings point towards a diversity of functional niches within root systems, which dynamically change in response to soil nutrients. Our study provides new insights for understanding root-microbiota interactions of individual root types to environmental stimuli aiming to improve plant growth and fitness.\nC_LI

microbiology

SFMetaDB: A Comprehensive Annotation of Mouse RNA Splicing Factor RNA-Seq Datasets

Although the number of RNA-Seq datasets deposited publicly has increased over the past few years, incomplete annotation of the associated metadata limits their potential use. Because of the importance of RNA splicing in diseases and biological processes, we constructed a database called SFMetaDB by curating datasets related with RNA splicing factors. Our effort focused on the RNA-Seq datasets in which splicing factors were knocked-down, knocked-out or over-expressed, leading to 75 datasets corresponding to 56 splicing factors. These datasets can be used in differential alternative splicing analysis for the identification of the potential targets of these splicing factors and other functional studies. Surprisingly, only [~]15% of all the splicing factors have been studied by loss- or gain-of-function experiments using RNA-Seq. In particular, splicing factors with domains from a few dominant Pfam domain families have not been studied. This suggests a significant gap that needs to be addressed to fully elucidate the splicing regulatory landscape. Indeed, there are already mouse models available for [~]20 of the unstudied splicing factors, and it can be a fruitful research direction to study these splicing factors in vitro and in vivo using RNA-Seq.\n\nDatabase URLhttp://sfmetadb.ece.tamu.edu/

bioinformatics

l1kdeconv: an R package for peak calling analysis with LINCS L1000 data

BackgroundLINCS L1000 is a high-throughput technology that allows gene expression measurement in a large number of assays. However, to fit the measurements of ~ 1000 genes in the ~ 500 color channels of LINCS L1000, every two landmark genes are designed to share a single channel. Thus, a deconvolution step is required to infer the expression values of each gene. Any errors in this step can be propagated adversely to the downstream analyses.\n\nResultsWe presented a LINCS L1000 data peak calling R package l1kdeconv based on a new outlier detection method and an aggregate Gaussian mixture model (AGMM). Upon the remove of outliers and the borrowing information among similar samples, l1kdeconv showed more stable and better performance than methods commonly used in LINCS L1000 data deconvolution.\n\nConclusionsBased on the benchmark using both simulated data and real data, the l1kdeconv package achieved more stable results than the commonly used LINCS L1000 data deconvolution methods.\n\nContactpengyu.bio@gmail.com

bioinformatics