Search bioRxiv⌕ Search

Biology subjects

McLoughlin, K. S.

Publications and source records attributed to McLoughlin, K. S..

2 recordsLinked to original sources

Generative Molecular Design and Experimental Validation of Selective Histamine H1 Inhibitors

Generative molecular design (GMD) is an increasingly popular strategy for drug discovery, using machine learning models to propose, evaluate and optimize chemical structures against a set of target design criteria. We present the ATOM-GMD platform, a scalable multiprocessing framework to optimize many parameters simultaneously over large populations of proposed molecules. ATOM-GMD uses a junction tree variational autoencoder mapping structures to latent vectors, along with a genetic algorithm operating on latent vector elements, to search a diverse molecular space for compounds that meet the design criteria. We used the ATOM-GMD framework in a lead optimization case study to develop potent and selective histamine H1 receptor antagonists. We synthesized 103 of the top scoring compounds and measured their properties experimentally. Six of the tested compounds bind H1 with Kis between 10 and 100 nM and are at least 100-fold selective relative to muscarinic M2 receptors, validating the effectiveness of our GMD approach.

pharmacology and toxicology↗

Model Choice Metrics to Optimize Profile-QSAR Performance

BackgroundPredicting molecular activity against protein targets is difficult because of the paucity of experimental data. Approaches like multitask modeling and collaborative filtering seek to improve model accuracy by leveraging results from multiple targets, but are limited because different compounds are measured with different assays, leading to sparse data matrices. Profile-QSAR (pQSAR) 2.0 addresses this problem by fitting a series of partial least squares models for each target, using as features the predictions from single-task models on the remaining targets. This method has been shown to produce better results than single task and multitask models. However, the factors determining the success of pQSAR 2.0 have as yet not been characterized. In this paper we examine the experimental conditions that lead to better pQSAR models. We limit the amount of data available to the method by retraining with decreasing amounts of data and explore the models ability to generalize to compounds that have never been assayed. Finally, we look at the properties of training data needed to demonstrate pQSAR improvement. ResultsWe apply pQSAR 2.0 on a collection of GPCR and safety targets collected from Drug Target Commons, ExcapeDB, and ChEMBL. We found that pQSAR improved models on 34 of the 149 assays selected. In the other 115 assays, single task random forests offered better performance. There are many factors that contribute to an increase in performance, but the main factor is compound assay coverage. The pQSAR model improves when more compounds are measured in multiple assays. ConclusionIt is necessary to consider the available data before applying pQSAR. Successful pQSAR models require a profile made of correlated targets that share compounds with other assays. This technique is best used when experimental data is available as random forest regressors often do not generalize well enough for virtual drug search applications.

bioinformatics↗