bioRxiv · 10.64898/2026.02.15.706011
MolDeBERTa: Foundational Model for Physicochemical and Structural-Informed Molecular Representation Learning
Abstract
Foundational models that learn the "language" of molecules are essential for accelerating material and drug discovery. These self-learning models can be trained on large collections of unlabelled molecules, enabling applications such as property prediction, molecule design, and screening for specific functions. However, existing molecular language models rely on masked language modeling, a generic token-level objective that is agnostic to physico-chemical and substructure molecular properties. Here we introduce MolDeBERTa, a chemistry-informed self-supervised molecular encoder built upon the DeBERTaV2 architecture with byte-level Byte-Pair Encoding (BPE) tokenization. MolDeBERTa is pre-trained on up to 123 million SMILES from PubChem using three novel pretraining objectives designed to inject strong inductive biases for molecular properties and substructure similarity directly into the latent space. The model is systematically investigated across three architectural scales, two dataset sizes, and five distinct pretraining objectives, of which three are novel and two are adapted from prior work. When evaluated on 9 MoleculeNet benchmarks, MolDeBERTa achieves the best overall performance on 4 out of 9 tasks and outperforms SMILES-based encoders on 7 out of 9 tasks, with up to a 16% reduction in regression error, and improvements of up to 2.2 ROC-AUC points on classification tasks. The source code, pretrained checkpoints, and datasets are publicly available at https://github.com/pcdslab/MolDeBERTa and https://huggingface.co/collections/SaeedLab/moldeberta. CCS ConceptsO_LIComputing methodologies [->] Neural networks; Natural language processing; Learning paradigms. C_LI
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
de Oliveira, G. B., Saeed, F.. 2026-02-17. MolDeBERTa: Foundational Model for Physicochemical and Structural-Informed Molecular Representation Learning. https://doi.org/10.64898/2026.02.15.706011
Cite the original work for its findings. Save a collection to share your selection of sources.