Search bioRxiv⌕ Search

Biology subjects

Daina, A.

Publications and source records attributed to Daina, A..

2 recordsLinked to original sources

Knowledge-guided machine-learning and reverse screening combined method to predict cancer cell line responses to cytotoxic molecules

Estimating the cell line targets of cytotoxic small molecules is important for drug discovery and central for targeted therapy in oncology. Accurate prediction of sensitive cell lines enables early identification of efficacy and toxicity, optimization of drug selectivity, and can foster drug repurposing. While most bioactive compounds interact with multiple macromolecular targets, the cytotoxicity encompasses diverse complex biological and chemical mechanisms that could even not all be related to binding to macromolecules, making the prediction of cytotoxicity specificity particularly challenging. To address early-phase prediction of cancer cell line targets of cytotoxic compounds, we developed a method combining a machine-learning classification model with a ligand-based reverse screening procedure able to rank cell-line from the most probable to the least probable target of any cytotoxic molecule. The development focused on addressing the challenges related to the scarcity of available experimental data on non-cytotoxic compounds. A knowledge-guided generation of realistic alleged inactives allowed to train several binary logistic regression models. The most robust classification model was trained on 164,134 cytotoxic compounds extracted from ChEMBL to generate a score of predicted sensitivity of cell lines for any new cytotoxic molecule. The method demonstrated strong predictive ability, recovering at least one experimental target within the 15 most probable cell-lines for 71% of nearly 11,000 external cytotoxic compounds tested across 1018 cancer cell lines.

bioinformatics↗

Rethinking Molecular Beauty in the Deep Learning Era

As (deep) generative chemistry rapidly enters the landscape of drug discovery, evaluating the models and their structural output relevance, tractability, and innovation potential becomes both a scientific and practical challenge. Existing metrics such as QED and SA Score, though widely adopted, are rooted in biased historical datasets and often fail to capture what medicinal chemists truly seek: compounds with meaningful pharmacological potential and realistic tractability. In this work, we introduce the OiiSTER-map--a simple, intuitive, and interpretable two-dimensional heatmap that classifies molecules based on bioactivity-informed usuality and structural elaboration derived from molecular fingerprints. Unlike traditional filters, the OiiSTER-map helps identify not only the drug discovery chemical "sweet spot". (Regular/Balanced), but also underappreciated territories such as Minimal/Unusual or Over-elaborated/Trivial regions, offering actionable insights into compound quality, relevance, and diversity. We hope that a bioactivity-informed, structurally aware, and easily interpretable tool like the OiiSTER-map, employed in combination with other well implemented metrics, can be decisive to go beyond current limitations of assessing (deep) generative models and to ensure a more mechanistically relevant, nuanced and useful evaluation of the thus designed virtual molecules.

bioinformatics↗