Search bioRxivSearch

Biology subjects

Pirvan, L.

Publications and source records attributed to Pirvan, L..

2 recordsLinked to original sources

Identifying regulatory and spatial genomic architectural elements using cell type independent machine and deep learning models

Chromosome conformation capture methods such as Hi-C enables mapping of genome-wide chromatin interactions and is a promising technology to understand the role of spatial chromatin organisation in gene regulation. However, the generation and analysis of these data sets at high resolutions remain technically challenging and costly. We developed a machine and deep learning approach to predict functionally important, highly interacting chromatin regions (HICR) and topologically associated domain (TAD) boundaries independent of Hi-C data in both normal physiological states and pathological conditions such as cancer. This approach utilises gradient boosted trees and convolutional neural networks trained on both Hi-C and histone modification epigenomic data from three different cell types. Given only epigenomic modification data these models are able to predict chromatin interactions and TAD boundaries with high accuracy. We demonstrate that our models are transferable across cell types, indicating that combinatorial histone mark signatures may be universal predictors for highly interacting chromatin regions and spatial chromatin architecture elements.

genomics

Pangaea: A modular and extensible collection of tools for mining context dependent gene relationships from the biomedical literature

MotivationPangaea is a scalable and extensible command line interface (CLI) software that integrates gene-relationship detection features to extract context-dependent structured gene-gene and gene-term relationships from the biomedical literature. It provides computational methods to identify biological relationships between a collection of genes and can be used to search and extract different types of contextual relationships amongst genes. ResultsWe implemented a CLI-based software for downloading PubMed articles and extracting gene relationships from abstracts using natural language processing methods. In terms of scalability, the software was designed to support the retrieval and processing of millions of articles whilst minimising memory requirements and optimising for parallel processing on multiple CPU cores. To allow extensibility, the tool permits the use of contextual custom-made models for the text processing parts, and the output is serialised as JSON objects to allow flexible post-processing workflows. AvailabilityThe software is available online at: https://github.com/ss-lab-cancerunit/pangaea

bioinformatics