Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.07.02.600897

Prediction of single-cell chromatin compartments from single-cell chromosome structures by MaxComp

Abstract

The genome is partitioned into distinct chromatin compartments with at least two main classes, a transcriptionally active A and an inactive B compartment, corresponding mostly to the segregation of euchromatin and heterochromatin. Chromatin within the same compartment has a higher tendency to interact with itself than with regions in opposing compartments. A/B compartments are traditionally derived from ensemble Hi-C contact matrices through principal component analysis of their covariance matrices. However, defining compartments in single cells from single-cell Hi-C maps is non trivial due to sparsity of the data and the fact that homologous copies are typically not resolved. Here we present an unsupervised approach, named MaxComp, to determine single-cell A/B compartments from geometric considerations in 3D chromosome structures, either from multiplexed FISH imaging or from models derived from Hi-C data. By representing each single-cell structure as an undirected graph with edge-weights encoding structural information, the problem of predicting chromosome compartments can be transformed to an alternative form of the Max-cut problem, a semidefinite graph programming method (SPD) to determine an optimal division of a chromosome structure graph into two structural compartments. Our results show that compartment annotations from principal component analysis of ensemble Hi-C data can be perfectly reproduced as population averages of our single-cell compartment predictions. We therefore prove that compartment predictions can be achieved from geometric considerations alone using 3D coordinates of chromatin regions together with information about their nuclear microenvironment. Our results reveal substantial cell-to-cell heterogeneity of compartments in a cell population, which substantially differs between individual genomic regions. Moreover, by applying our approach to multiplexed FISH tracing experiments, our method sheds light on the relationship between single-cell compartment annotations and gene transcriptional activity in single cells. Overall our approach provides new insights into single-cell chromatin condensation, relationship between population and single-cell chromatin compartmentalization, the cell-to-cell variations of chromatin compartments and its impact on gene transcription. Author SummaryChromosome conformation capture and imaging techniques revealed the segregation of genomic chromatin into at least two functional compartments. Hi-C contact frequency matrices show checkerboard-like patterns indicating that chromatin regions are divided into at least two states, possibly a result of phase separation. Chromatin regions in the same state have preferential interactions with each other, often over extended sequence distances, while interactions to regions in the opposing state are minimized. Principal component analysis (PCA) on ensemble Hi-C contact frequency matrices can identify these compartment states. However, because the compartment annotations are derived from a cell population, this method cannot provide information about compartments in single cells. Here in this study, we introduce an unsupervised method to predict single-cell compartments using graph-based programming, which utilizes only structural information in single cells. Our results demonstrate that PCA-based ensemble compartment annotations can be reproduced as population averages of our single-cell compartment predictions. Moreover, our results reveal the cell-to-cell heterogeneity of compartments in a cell population, which shows significant disparities among different chromatin regions. Moreover, by applying our approach to multiplexed FISH tracing experiments, our method reveals the relationship between single-cell compartment annotations and gene transcriptional activity in single cells. Finally, our approach also allows us to relate chromatin structural features in single cells with compartment properties. Comparison with other existing approaches showed that our method produces overall better compartmentalization scores in single cells.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhan, Y., Musella, F., Alber, F.. 2024-07-04. Prediction of single-cell chromatin compartments from single-cell chromosome structures by MaxComp. https://doi.org/10.1101/2024.07.02.600897

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

spatialMET: an open and scalable framework for spatial metabolomics analysis

Mass spectrometry imaging (MSI) enables spatially resolved metabolomics in intact tissue sections, but analysis remains challenging at scale. Existing MSI workflows often require users to combine multiple software tools, while others rely on proprietary vendor software that limits interoperability and reproducibility. To address these challenges, we developed spatialMET, an open-source framework that provides an end-to-end workflow for MSI analysis. spatialMET provides a unified platform for preprocessing, spatial domain detection, and visualization. Downstream analyses include differential abundance testing, spatial autocorrelation and gradient analysis, dimensionality reduction, and correlation network analysis. Spatial domain detection uses hcdist, a C-based hierarchical clustering implementation that substantially reduces runtime and memory use relative to existing R-based approaches. spatialMET can be run through an interactive R Shiny application or as a standalone command-line workflow for larger datasets or high-performance computing environments. Applied to mouse small cell lung cancer MALDI-MSI data containing 284,673 pixels, spatialMET identified tumor-associated, stromal, and adjacent lung spatial domains that aligned with matched histology. Differential abundance analysis identified 117 m/z features that differed between tumor and stromal regions, while spatial autocorrelation analyses revealed spatially structured abundance patterns. Applying spatialMET to mouse lung adenocarcinoma data from an entire lung lobe containing 338,477 pixels further demonstrated scalability and captured spatial heterogeneity across tumor and surrounding lung tissue. In summary, spatialMET provides a scalable, open-source framework for end-to-end spatial metabolomics analysis, and it is distributed as a Docker container for reproducible deployment. Source code and installation instructions are available at https://github.com/biodatalab/spatialMET.

bioinformatics↗

Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach

Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.

bioinformatics↗

An Information Geometry approach to model topological trajectories and Gene Expression Radius from UMAP geometry.

Understanding the relationship between gene expression dynamics and cellular identity remains a central challenge in single cell biology. Here, we introduce a novel computational and mathematical framework that integrates information geometry, fuzzy topology, and UMAP analysis to model gene expression landscapes derived from single cell RNA sequencing data. We formalize gene expression data as a fuzzy topological space, where interactions between expression points are governed by probabilistic distributions inspired by manifold learning approaches such as UMAP. Within this framework, we define an information geometric structure through a Fisher metric induced by these distributions, enabling the computation of geodesic trajectories that capture cellular differentiation processes. A key contribution of this work is the derivation of analytical conditions, expressed as expression radius formulas, that characterize local neighborhoods in gene expression space. These conditions allow for the identification of genes associated with stem cell states and predictions in transitional cell types in future work. Application of the proposed framework to single cell datasets reveals biologically meaningful gene sets enriched in key regulatory pathways and transcription factors, demonstrating the capacity of our approach to uncover latent structure in complex gene expression data. Our results suggest that integrating differential geometry with statistical learning theory offers a powerful paradigm for modeling genotype and phenotype relationships and cellular state transitions, with potential implications for precision medicine and systems biology.

bioinformatics↗