Search bioRxiv⌕ Search

Biology subjects

Zsigmond, K.

Publications and source records attributed to Zsigmond, K..

4 recordsLinked to original sources

The Shape of Chemical Space

The concept of chemical space is critical in cheminformatics, medicinal chemistry, and machine learning applications. Despite this, the high dimensionality of molecular representations greatly complicates its sampling, analysis, and visualization. A popular approach to overcome problem is to project these representations to a "human-manageable" subspace, usually containing only two dimensions. Non-linear dimensionality reduction techniques are by far the preferred strategy, following the reasoning that their flexibility can accommodate any arbitrary distribution originally present in the high-dimensional space. However, this ignores the elevated computational cost of these methods and the difficulty in tuning their hyper-parameters. Here, we show that basic properties of the metrics used in the original space can be used to infer the shape of the chemical space, which in turns suggests an optimal strategy to project chemical information to lower dimensions. The key insight is to realize that, no matter the set of molecules, their fingerprint representation can be considered to lie on a hyper-spherical surface. The smooth nature of this manifold means that we can use clustering to identify locally-dense sectors of chemical space, and selectively project them simply using linear (hyper-parameter free) methods, like principal component analysis. This approach surpasses non-linear techniques in several neighborhood preservation metrics, while only requiring a fraction of the computational cost. This pipeline is implemented in our N-Ary Mapping Interface (NAMI: https://github.com/mqcomplab/NAMI), which we tested in the visualization of 10 million molecules.

bioinformatics↗

Undersampling techniques for non-linear chemical space visualization

The visualization of high-dimensional chemical space is a critical tool for understanding molecular diversity, structure-property relationships, and for guiding compound selection. However, the performance of non-linear dimensionality reduction (DR) techniques like t-Stochastic Neighborhood Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), and Generative Topographic Mapping (GTM) are often susceptible to the choice of hyperparameters, along with the high cost of their training for large datasets. In this study, we investigated the effect of undersampling methods on the choice of hyperparameter selection for these non-linear dimensionality reduction methods. Our results demonstrate that selecting small representative subsets of chemical data not only reduces computational costs associated with hyperparameter training but also serves as an innovative means to train non-linear DR methods, leading to projections that better preserve the local structure within the chemical space.

bioinformatics↗

Is Tanimoto a metric?

No. However, here we show how to generate a metric consistent with the Tanimoto similarity. We also explore new properties of this index, and how it relates to other popular alternatives.

bioinformatics↗

SHINE: Deterministic Many-to-Many clustering of Molecular Pathways

State-of-the-art molecular dynamics (MD) simulation methods can generate diverse ensembles of pathways for complex biological processes. Analyzing these pathways using statistical mechanics tools demands identifying key states that contribute to both the dynamic and equilibrium properties of the system. This task becomes especially challenging when analyzing multiple MD simulations simultaneously, a common scenario in enhanced sampling techniques like the weighted ensemble strategy. Here, we present a new module of the MDANCE package designed to streamline the analysis of pathway ensembles. This module integrates n-ary similarity, cheminformatics-inspired tools, and hierarchical clustering to improve analysis efficiency. We present the theoretical foundation behind this approach, termed Sampling Hierarchical Intrinsic N-ary Ensembles (SHINE), and demonstrate its application to simulations of alanine dipeptide and adenylate kinase.

biophysics↗