bioRxiv · 10.64898/2026.09.16.751426
Machine Learning for Toxicity Prediction in Low-Sample Molecular Classes
Abstract
Deep learning models such as Chemprop have advanced quantitative molecular property prediction, but their reliance on large training sets limits use in data-scarce domains. We propose a framework that fine-tunes a general baseline model trained on publicly available data on small, class-specific datasets. The resulting models retain the baseline's generalization ability while gaining class-specific accuracy and produce probabilistic outputs that capture uncertainty in the training data. We demonstrate the approach on three toxicity classes defined by a common core structure, target, or mode of action: (i) organophosphates, (ii) androgen receptor antagonists, and (iii) estrogen receptor beta antagonists. Each fine-tuned model outperforms classical machine-learning methods and the EPA TEST tool. The probabilistic nature of the predictions enables prioritization of compounds for experimental validation and seamless integration with data streams of varying quality, supporting iterative decision-making in chemical safety and drug discovery.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Barajas, C., Dunphy, L., Mullany, L., Tiburzi, O., Lloyd, E.. 2026-09-18. Machine Learning for Toxicity Prediction in Low-Sample Molecular Classes. https://doi.org/10.64898/2026.09.16.751426
Cite the original work for its findings. Save a collection to share your selection of sources.