Search bioRxiv⌕ Search

bioRxiv · 10.64898/2026.05.20.726468

MolCodon: A Codon-Based Molecular Language for InterpretableStructural Representation and Similarity Search

Abstract

Molecular representation determines which aspects of chemical structure can be learned, compared, and interpreted in computational drug discovery. Existing encodings typically emphasize either compact string description, as in SMILES and SELFIES, or efficient similarity search, as in circular fingerprints, but they may not simultaneously provide deterministic sequence structure, graph-level interpretability, pharmacophore annotation, and high-fidelity molecular reconstruction. Here, we introduce MolCodon, a codon-based molecular language that represents small molecules as deterministic sequences of fixed-width three-character tokens over a five-symbol alphabet, C, N, O, S, and X. Inspired by the triplet organization of the genetic code, MolCodon assigns chemically defined codon families to atoms, bonds, ring and branch topology, fused-ring references, pharmacophore features, bond mobility, charge, and stereochemistry. A deterministic graph traversal with ring-contiguity preservation produces sequences in which chemically meaningful substructures remain locally organized and traceable to the underlying molecular graph. Across around 2,9 million molecules from six commercial screening libraries, MolCodon achieved 98.93% InChIKey-level round-trip fidelity, supporting its use as a high-fidelity sequence representation for drug-like chemistry. MolCodon-derived sparse sequence and trace features further outperformed SELFIES and Group SELFIES across ten QSAR tasks and exceeded classical fingerprint baselines in six out of ten tasks. As an application of the representation, MolCodon BLAST similarity engine decomposes molecular similarity into ring topology, branch context, attachment architecture, and pharmacophore correspondence, enabling interpretable scaffold-hopping searches. In a PARP1 virtual screening study, MolCodon retrieved scaffold-diverse candidates to a known PARP-1 inhibitor Olaparib. Together, these results establish MolCodon as a new molecular representation paradigm that transforms chemical graphs into high-fidelity, interpretable, and alignment-compatible codon sequences, opening a direct path for bioinformatics-inspired analysis of small-molecule chemical space. The MolCodon encoder, decoder, and BLAST similarity engine are freely available as open-source software at https://github.com/DurdagiLab/MolCodon

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sayyah, E., Kurul, E., Tunc, H., DURDAGI, S.. 2026-05-21. MolCodon: A Codon-Based Molecular Language for InterpretableStructural Representation and Similarity Search. https://doi.org/10.64898/2026.05.20.726468

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

QuickSeg: A fast, versatile and accurate algorithm for genomic copy number segmentation using dynamic programming

Copy number alterations are among the most common genomic aberrations in cancer and their accurate identification relies on robust segmentation of sequencing read-depth signals. Existing segmentation methods typically balance computational efficiency against segmentation accuracy and remain sensitive to technical artifacts present in sequencing data. Here, we present QuickSeg, a fast and versatile methodology that uses an exact dynamic programming algorithm to detect copy number segments using median-based error function. Motivated by the observation that sequencing depth distributions contain a small but pervasive population of outlying observations, this approach provides increased robustness to technical noise while simultaneously reducing the computational complexity of the segmentation problem. Across whole-genome sequencing of cancer cohorts, using breakpoint-supported somatic copy number alterations, we demonstrate improved segmentation precision over two widely used baseline methods, Circular Binary Segmentation (CBS) and Piecewise Constant Fitting (PCF), across a broad range of sensitivity thresholds. QuickSeg also consistently outperformed both methods with respect to runtime and memory usage. Collectively, our results show that robust median-based optimization provides both biological and computational advantages for copy number segmentation, enabling accurate analysis of large sequencing cohorts with minimal computational requirements.

bioinformatics↗

AltraFlowSOM: A Semi-Supervised Framework for Imaging Mass Cytometry Phenotyping

Imaging Mass Cytometry (IMC) enables the simultaneous quantification of 40+ protein markers at single cell resolution in tissue, however biologically faithful phenotyping at scale remains a critical bottleneck. Unsupervised clustering fragments coherent populations or conversely merges biologically incoherent ones into a single cluster, supervised classifiers impose a closed vocabulary, and the presence of rare subsets (encoding clinically relevant biology) in conjunction with abundant subsets may be detrimental to detection performances. We present AltraFlowSOM, a semi-supervised extension of FlowSOM that embeds partial expert annotations directly into self-organizing map training via a two-layer SuperSOM architecture, balancing label-guided topology anchoring with unsupervised discovery. By anchoring the map to biologically labelled reference points, AltraFlowSOM circumvents the canonical dependency between batch correction and clustering. Evaluated under Leave-one-out cross validation on two independent IMC cohorts, Lupus Nephritis (n=22 ROIs) and Sjogren syndrome (n=10 ROIs), AltraFlowSOM outperformed all unsupervised and supervised baseline on Adjusted Rand Index, F1 scores (macro and weighted), weighted purity and in the identification of rare populations. The median Treg cell recovery exceeded that of all comparator methods. AltraFlowSOM resolves the scalability-alignment-discovery trilemma, by establishing a semi-supervised SOM as a generalizable method for high dimensional IMC phenotyping.

bioinformatics↗

Modelling interpretable patient-level representationsfrom structured and simple multimodal data

Patient cohort profiling increasingly includes structured views for multiple modalities, such as single-cell RNA sequencing, spatial transcriptomics or proteomics, and histology, each providing multiple subobservations per patient, including single cells, spatial spots or patches. To model such data along with simple patient-level views, current multimodal integration methods typically rely on separately precomputed summaries and fail to fully leverage information in structured views. Here we present FACTMx, a variational framework that jointly models structured and simple views to learn interpretable patient-level representations. FACTMx couples latent patient factors with subobservation clustering and per-patient component proportions, enabling direct interpretation and downstream association analyses. The framework supports different structured-view mixture assumptions, including topic- and Gaussian-structured data, while retaining modular encoder-decoder parameterisations. In simulations spanning sparse and dense dependencies and multiple noise regimes, FACTMx improved reconstruction, integration and recovery of structured components relative to previous methods. Applied to non-small cell lung cancer cohorts, FACTMx captured survival-associated latent signals linked to immune microenvironments, gene expression pathways and spatially coherent histological patterns. In a longitudinal coronary syndrome cohort, FACTMx highlighted an outcome-associated axis connected to ejection-fraction change, immune cell states, soluble mediators and cardiac injury markers. These results support joint structured-simple modelling for interpretable multimodal patient stratification.

bioinformatics↗