Search bioRxiv⌕ Search

Biology subjects

Korsak, S.

Publications and source records attributed to Korsak, S..

5 recordsLinked to original sources

The challenge of chromatin model comparison and validation a project from the first international 4D Nucleome Hackathon

The computational modeling of chromatin structure is highly complex and challenging due to the hierarchical organization of chromatin, which reflects its diverse biophysical principles, as well as inherent dynamism, which underlies its complexity. The variety of methods for chromatin structure modeling, which are based on different approaches, assumptions and scales of modeling, suggests that there is a necessity for a comprehensive benchmark. This inspired us to conduct a project at the NIH-funded 4D Nucleome Hackathon on March 18-21, 2024 at The University of Washington in Seattle, USA. The hackathon provided an amazing opportunity to gather an international, multi-institutional and unbiased group of experts to discuss, understand and undertake the challenges of chromatin model comparison and validation. These challenges seem straightforward in theory, however in practice, they are challenging and ambiguous. To address them, we developed a bioinformatics workflow for chromatin model comparison and validation, in which we use distance matrices to represent chromatin models, and we calculate Spearman correlation coefficients between pairs of matrices to estimate correlations between models, as well as between models and experimental data. During the 4-day hackathon, we tested our workflow on several distinct software packages for chromatin structure modeling and we discovered several challenges that include: 1) different aspects of chromatin biophysics and scales complicate model comparisons, 2) expertise in biology, bioinformatics, and physics is necessary to conduct a comprehensive research on chromatin structure, 3) bioinformatic software, which is often developed in academic settings, is characterized by insufficient support and documentation. Therefore, our work constitutes a way to advance the modeling of the 3D organization of the human genome, while emphasizing the importance of establishing guidelines for software development and standardization.

genomics↗

Multiscale Molecular Modelling of Chromatin with MultiMM: From Nucleosomes to the Whole Genome

MotivationWe present a user-friendly 3D chromatin simulation model based on OpenMM, addressing the challenges posed by existing models with use-specific implementations. Our approach employs a multi-scale energy minimization strategy, capturing chromatins hierarchical structure. Initiating with a Hilbert curve-based structure, users can input files specifying nucleosome positioning, loops, compartments, or subcompartments. ResultsThe model utilizes an energy minimization approach with a large choice of numerical integrators, providing the entire genomes structure within minutes. Output files include the generated structures for each chromosome, offering a versatile and accessible tool for chromatin simulation in bioinformatics studies. Furthermore, MultiMM is capable of producing nucleosomeresolution structures by making simplistic geometric assumptions about the structure and the density of nucleosomes on the DNA. Code availabilityOpen-source software and the manual are freely available on https://github.com/SFGLab/MultiMM.

bioinformatics↗

Improved cohesin HiChIP protocol and bioinformatic analysis for robust detection of chromatin loops and stripes

Chromosome Conformation Capture (3C) methods, including Hi-C (a high-throughput variation of 3C), detect pairwise interactions between DNA regions, enabling the reconstruction of chromatin architecture in the nucleus. HiChIP is a modification of the Hi-C experiment, which includes a chromatin immunoprecipitation step (ChIP), allowing genome-wide identification of chromatin contacts mediated by a protein of interest. In mammalian cells, cohesin protein complex is one of the major players in the establishment of chromatin loops. We present an improved cohesin HiChIP experimental protocol. Using comprehensive bioinformatic analysis, we show that performing cohesin HiChIP with two cross-linking agents (formaldehyde [FA] and EGS) instead of the typically used FA alone, results in a substantially better signal-to-noise ratio, higher ChIP efficiency and improved detection of chromatin loops and architectural stripes. Additionally, we propose an automated pipeline called nf-HiChIP (https://github.com/SFGLab/hichip-nf-pipeline) for processing HiChIP samples starting from raw sequencing reads data and ending with a set of significant chromatin interactions (loops), which allows efficient and timely analysis of multiple samples in parallel, without the need of additional ChIP-seq experiments. Finally, using novel approaches for biophysical modelling and stripe calling we generate accurate loop extrusion polymer models for a region of interest and a detailed picture of architectural stripes, respectively.

genomics↗

LoopSage: An Energy-Based Monte Carlo approach for the Loop Extrusion Modelling of Chromatin

The connection between the patterns observed in 3C-type experiments and the modeling of polymers remains unresolved. This paper presents a simulation pipeline that generates thermodynamic ensembles of 3D structures for topologically associated domain (TAD) regions by loop extrusion model (LEM). The simulations consist of two main components: a stochastic simulation phase, employing a Monte Carlo approach to simulate the binding positions of cohesins, and a dynamical simulation phase, utilizing these cohesins positions to create 3D structures. In this approach, the systems total energy is the combined result of the Monte Carlo energy and the molecular simulation energy, which are iteratively updated. The structural maintenance of chromosomes (SMC) protein complexes are represented as loop extruders, while the CCCTC-binding factor (CTCF) locations on DNA sequence are modeled as energy minima on the Monte Carlo energy landscape. Finally, the spatial distances between DNA segments from ChIA-PET experiments are compared with the computer simulations, and we observe significant Pearson correlations between predictions and the real data. LoopSage model offers a fresh perspective on chromatin loop dynamics, allowing us to observe phase transition between sparse and condensed states in chromatin.

genomics↗

Multi-scale phase separation by explosive percolation with single chromatin loop resolution

The 2m-long human DNA is tightly intertwined into the cell nucleus of the size of 10m. The DNA packing is explained by folding of chromatin fiber. This folding leads to the formation of such hierarchical structures as: chromosomal territories, compartments; densely packed genomic regions known as Chromatin Contact Domains (CCDs), and loops. We propose models of dynamical genome folding into hierarchical components in human lymphoblastoid, stem cell, and fibroblast cell lines. Our models are based on explosive percolation theory. The chromosomes are modeled as graphs where CTCF chromatin loops are represented as edges. The folding trajectory is simulated by gradually introducing loops to the graph following various edge addition strategies that are based on topological network properties, chromatin loop frequencies, compartmentalization, or epigenomic features. Finally, we propose the genome folding model - a biophysical pseudo-time process guided by a single scalar order parameter. The parameter is calculated by Linear Discriminant Analysis. We simulate the loop formation by using Loop Extrusion Model (LEM) while adding them to the system. The chromatin phase separation, where fiber folds into topological domains and compartments, is observed when the critical number of contacts is reached. We also observe that 80% of the loops are needed for chromatin fiber to condense in 3D space, and this is constant through various cell lines. Overall, our in-silico model integrates the high-throughput 3D genome interaction experimental data with the novel theoretical concept of phase separation, which allows us to model event-based time dynamics of chromatin loop formation and folding trajectories.

genomics↗