bioRxiv · 10.1101/150896
GEOracle: Mining perturbation experiments using free text metadata in Gene Expression Omnibus
Abstract
There exists over 2.5 million publicly available gene expression samples across 101,000 data series in NCBIs Gene Expression Omnibus (GEO) database. Due to the lack of the use of standardised ontology terms in GEOs free text metadata to annotate the experimental type and sample type, this database remains di[ffi]cult to harness computationally without significant manual intervention.\n\nIn this work, we present an interactive R/Shiny tool called GEOracle that utilises text mining and machine learning techniques to automatically identify perturbation experiments, group treatment and control samples and perform differential expression. We present applications of GEOracle to discover conserved signalling pathway target genes and identify an organ specific gene regulatory network.\n\nGEOracle is effective in discovering perturbation gene targets in GEO by harnessing its free text metadata. Its effectiveness and applicability has been demonstrated by cross validation and two real-life case studies. It opens up new avenues to unlock the gene regulatory information embedded inside large biological databases such as GEO. GEOracle is available at https://github.com/VCCRI/GEOracle.
Source connections
Explore related subjects
Keep this discovery
Djordjevic, D., Chen, Y. X., Kwan, S. L. S., Ling, R. W. K., Qian, G., Woo, C. Y. Y., Ellis, S. J., Ho, J. W. K.. 2017-06-16. GEOracle: Mining perturbation experiments using free text metadata in Gene Expression Omnibus. https://doi.org/10.1101/150896
Cite the original work for its findings. Save a collection to share your selection of sources.