Search bioRxiv⌕ Search

Biology subjects

Turon, G.

Publications and source records attributed to Turon, G..

2 recordsLinked to original sources

Machine learning approaches identify chemical features for stage-specific antimalarial compounds

Efficacy data from diverse chemical libraries, screened against the various stages of the malaria parasite Plasmodium falciparum, including asexual blood stage (ABS) parasites and transmissible gametocytes, serves as a valuable reservoir of information on the chemical space of compounds that are either active (or not) against the parasite. We postulated that this data can be mined to define chemical features associated with sole ABS activity and/or those that provide additional life cycle activity profiles like gametocytocidal activity. Additionally, this information could provide chemical features associated with inactive compounds, which could eliminate any future unnecessary screening of similar chemical analogues. Therefore, we aimed to use machine learning to identify the chemical space associated with stage-specific antimalarial activity. We collected data from various chemical libraries that were screened against the asexual (126 374 compounds) and sexual (gametocyte) stages of the parasite (93 941 compounds), calculated the compounds molecular fingerprints and trained machine learning models to recognize stage-specific active and inactiv compounds. We were able to build several models that predicts compound activity against ABS and dual-activity against ABS and gametocytes, with Support Vector Machines (SVM) showing superior abilities with high recall (90% and 66%) and low false positive predictions (15% and 1%). This allowed identification of chemical features enriched in active and inactive populations, an important outcome that could be mined for essential chemical features to streamline hit-to-lead optimization strategies of antimalarial candidates. The predictive capabilities of the models held true in diverse chemical spaces, indicating that the ML models are therefore robust and can serve as a prioritization tool to drive and guide phenotypic screening and medicinal chemistry programs. For Table of Contents Graphic Only O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=73 SRC="FIGDIR/small/553339v1_ufig1.gif" ALT="Figure 1"> View larger version (17K): org.highwire.dtl.DTLVardef@6c03d2org.highwire.dtl.DTLVardef@16eeefborg.highwire.dtl.DTLVardef@bcd95org.highwire.dtl.DTLVardef@e60be4_HPS_FORMAT_FIGEXP M_FIG C_FIG

biochemistry↗

First fully-automated AI/ML virtual screening cascade implemented at a drug discovery centre in Africa

We present ZairaChem, an artificial intelligence (AI)- and machine learning (ML)-based tool to train small-molecule activity prediction models. ZairaChem is fully automated, requires low computational resources and works across a broad spectrum of datasets, ranging from whole-cell growth inhibition assays to drug metabolism properties. The tool has been implemented end-to-end at the Holistic Drug Discovery and Development (H3D) Centre, the leading integrated drug discovery unit in Africa, at which no prior AI/ML capabilities were available. We have exploited in-house data collected from over a decade of drug discovery research in malaria and tuberculosis and built models to predict the outcomes of 15 key checkpoint assays. We subsequently deployed these models as a virtual screening cascade at an organisational scale to increase the hit rate of current experimental assays. We show how computational profiling of compounds, prior to synthesis and experimental testing, can increase the rate of progression by up to 40%. Moreover, we demonstrate that the approach can be applied to prioritise small molecules within a chemical series and to assess the likelihood of success of novel chemotypes, promoting efficient usage of limited experimental resources. This project is part of a first-of-its-kind collaboration between the H3D Centre, a research centre operating in a low-resource setting, and the Ersilia Open Source Initiative, a young tech non-profit devoted to building data science capacity in the Global South.

bioinformatics↗