Search bioRxiv⌕ Search

Biology subjects

Hernandez-Hernandez, S.

Publications and source records attributed to Hernandez-Hernandez, S..

3 recordsLinked to original sources

Machine Learning-Based Drug Response Prediction Identifies Novel Therapeutic Candidates for Colorectal Cancer Cell Line KM-12

One major hurdle in the discovery of new drugs is the limited capacity of traditional screening methods to efficiently explore the immense space of drug candidates. To bypass this difficulty, computational approaches such as virtual screening (VS) have been developed supported by Machine Learning (ML) models to predict the activities of drug-like molecules on a given target. We developed a ML model for VS of a library of 25 million highly-diverse synthesis-on-demand molecules against the KM-12 colorectal cancer (CRC) cell line. By combining different biological readouts, we discovered compounds with a novel chemical scaffold, not yet described in the context of CRC, with either a prominent cytotoxic or cytostatic effect. This finding underpins the strength of our ML-guided VS protocol in discovering promising CRC drug leads triggering altered biological responses.

cancer biology↗

Graph neural networks best guide phenotypic virtual screening on cancer cell lines

Artificial intelligence is increasingly driving early drug design, offering novel approaches to virtual screening. Phenotypic virtual screening (PVS) aims to predict how cancer cell lines respond to different compounds by focusing on observable characteristics rather than specific molecular targets. Some studies have suggested that deep learning may not be the best approach for PVS. However, these studies are limited by the small number of tested molecules as well as not employing suitable performance metrics and dissimilar-molecules splits better mimicking the challenging chemical diversity of real-world screening libraries. Here we prepared 60 datasets, each containing approximately 30,000 to 50000 molecules tested for their growth inhibitory activities on one of the NCI-60 cancer cell lines. We evaluated the performance of five machine learning algorithms for PVS on these 60 problem instances. To provide a comprehensive evaluation, we employed two model validation types: the random split and the dissimilar-molecules split. The models were primarily evaluated using hit rate, a more suitable metric in VS contexts. The results show that all models are more challenged by test molecules that are substantially different from those in the training data. In both validation types, the D-MPNN algorithm, a graph-based deep neural network, was found to be the most suitable for building predictive models for this PVS problem.

bioinformatics↗

Conformal prediction of molecule-induced cancer cell growth inhibition challenged by strong distribution shifts

The drug discovery process often employs phenotypic and target-based virtual screening to identify potential drug candidates. Despite the longstanding dominance of target-based approaches, phenotypic virtual screening is undergoing a resurgence due to its potential being now better understood. In the context of cancer cell lines, a well-established experimental system for phenotypic screens, molecules are tested to identify their whole-cell activity, as summarized by their half-maximal inhibitory concentrations. Machine learning has emerged as a potent tool for computationally guiding such screens, yet important research gaps persist, including generalization and uncertainty quantification. To address this, we leverage a clustering-based validation approach, called Leave Dissimilar Molecules Out (LDMO). This strategy enables a more rigorous assessment of model generalization to structurally novel compounds. This study focuses on applying Conformal Prediction (CP), a model-agnostic framework, to predict the activities of novel molecules on specific cancer cell lines. A total of 4320 independent models were evaluated across 60 cell lines, 5 CP variants, 2 set features, and training-test splits, providing strong and consistent results. From this comprehensive evaluation, we concluded that, regardless of the cell line or model, novel molecules with smaller CP-calculated confidence intervals tend to have smaller predicted errors once measured activities are revealed. It was also possible to anticipate the activities of dissimilar test molecules across 50 or more cell lines. These outcomes demonstrate the robust efficacy that LDMO-based models can achieve in realistic and challenging scenarios, thereby providing valuable insights for enhancing decision-making processes in drug discovery.

bioinformatics↗