Search bioRxiv⌕ Search

Biology subjects

Chilinski, M.

Publications and source records attributed to Chilinski, M..

5 recordsLinked to original sources

ConsensuSV-ONT - a modern method for accurate structural variant calling

The improvements in sequencing technology make the development of new tools for the detection of structural variance more and more common. However, the tools available for the long-read Oxford Nanopore sequencing are limited, and it is hard to choose one, which is the best. That is why there is a need to create a tool based on consensus that combines existing work in order to discover a set of high-quality, reliable structural variants that can be used for further downstream analysis. The field has also been subject to revolution in machine learning techniques, especially deep learning. In the spirit of the aforementioned need and developments, we propose a novel, fully automated ConsensuSV-ONT algorithm. The method uses six independent, state-of-the-art structural variant callers for long-read sequencing along with a convolutional neural network for filtering high-quality variants. We provide a runtime environment in the form of a docker image, wrapping a nextflow pipeline for efficient processing using parallel computing. The solution is complete in its form and is ready to use not only by computer scientists but accessible and easy to use for everyone working with Oxford Nanopore long-read sequencing data.

bioinformatics↗

Improved cohesin HiChIP protocol and bioinformatic analysis for robust detection of chromatin loops and stripes

Chromosome Conformation Capture (3C) methods, including Hi-C (a high-throughput variation of 3C), detect pairwise interactions between DNA regions, enabling the reconstruction of chromatin architecture in the nucleus. HiChIP is a modification of the Hi-C experiment, which includes a chromatin immunoprecipitation step (ChIP), allowing genome-wide identification of chromatin contacts mediated by a protein of interest. In mammalian cells, cohesin protein complex is one of the major players in the establishment of chromatin loops. We present an improved cohesin HiChIP experimental protocol. Using comprehensive bioinformatic analysis, we show that performing cohesin HiChIP with two cross-linking agents (formaldehyde [FA] and EGS) instead of the typically used FA alone, results in a substantially better signal-to-noise ratio, higher ChIP efficiency and improved detection of chromatin loops and architectural stripes. Additionally, we propose an automated pipeline called nf-HiChIP (https://github.com/SFGLab/hichip-nf-pipeline) for processing HiChIP samples starting from raw sequencing reads data and ending with a set of significant chromatin interactions (loops), which allows efficient and timely analysis of multiple samples in parallel, without the need of additional ChIP-seq experiments. Finally, using novel approaches for biophysical modelling and stripe calling we generate accurate loop extrusion polymer models for a region of interest and a detailed picture of architectural stripes, respectively.

genomics↗

HiCDiffusion - diffusion-enhanced, transformer-based prediction of chromatin interactions from DNA sequences

Prediction of chromatin interactions from DNA sequence has been a significant research challenge in the last couple of years. Several solutions have been proposed, most of which are based on encoder-decoder architecture, where 1D sequence is convoluted, encoded into the latent representation, and then decoded using 2D convolutions into the Hi-C pairwise chromatin spatial proximity matrix. Those methods, while obtaining high correlation scores and improved metrics, produce Hi-C matrices that are artificial - they are blurred due to the deep learning model architecture. In our study, we propose the HiCDiffusion model that addresses this problem. We first train the encoder-decoder neural network and then use it as a component of the diffusion model - where we guide the diffusion using a latent representation of the sequence, as well as the final output from the encoder-decoder. That way, we obtain the high-resolution Hi-C matrices that not only better resemble the experimental results - improving the Frechet inception distance by an average of 12 times, with the highest improvement of 35 times - but also obtain similar classic metrics to current state-of-the-art encoder-decoder architectures used for the task.

genomics↗

Enhanced performance of gene expression predictive models with protein-mediated spatial chromatin interactions

There have been multiple attempts to predict the expression of the genes based on the sequence, epigenetics, and various other factors. To improve those predictions, we have decided to investigate adding protein-specific 3D interactions that play a major role in the compensation of the chromatin structure in the cell nucleus. To achieve this, we have used the architecture of one of the state-of-the-art algorithms, ExPecto (J. Zhou et al., 2018), and investigated the changes in the model metrics upon adding the spatially relevant data. We have used ChIA-PET interactions that are mediated by cohesin (24 cell lines), CTCF (4 cell lines), and RNAPOL2 (4 cell lines). As the output of the study, we have developed the Spatial Gene Expression (SpEx) algorithm that shows statistically significant improvements in most cell lines.

genomics↗

Multi-scale phase separation by explosive percolation with single chromatin loop resolution

The 2m-long human DNA is tightly intertwined into the cell nucleus of the size of 10m. The DNA packing is explained by folding of chromatin fiber. This folding leads to the formation of such hierarchical structures as: chromosomal territories, compartments; densely packed genomic regions known as Chromatin Contact Domains (CCDs), and loops. We propose models of dynamical genome folding into hierarchical components in human lymphoblastoid, stem cell, and fibroblast cell lines. Our models are based on explosive percolation theory. The chromosomes are modeled as graphs where CTCF chromatin loops are represented as edges. The folding trajectory is simulated by gradually introducing loops to the graph following various edge addition strategies that are based on topological network properties, chromatin loop frequencies, compartmentalization, or epigenomic features. Finally, we propose the genome folding model - a biophysical pseudo-time process guided by a single scalar order parameter. The parameter is calculated by Linear Discriminant Analysis. We simulate the loop formation by using Loop Extrusion Model (LEM) while adding them to the system. The chromatin phase separation, where fiber folds into topological domains and compartments, is observed when the critical number of contacts is reached. We also observe that 80% of the loops are needed for chromatin fiber to condense in 3D space, and this is constant through various cell lines. Overall, our in-silico model integrates the high-throughput 3D genome interaction experimental data with the novel theoretical concept of phase separation, which allows us to model event-based time dynamics of chromatin loop formation and folding trajectories.

genomics↗