Search bioRxiv⌕ Search

Biology subjects

Mokaya, M.

Publications and source records attributed to Mokaya, M..

2 recordsLinked to original sources

Expanding the Synthome: Generating and integrating novel reactions into retrosynthesis tools.

Synthesising compounds is a major component of the time and cost of drug discovery. Retrosynthesis approaches have grown in prominence in efficiently predicting synthetic routes. For such tools to be truly impactful they should be able to effectively incorporate rarely seen chemical reactions and synthesise unknown molecules. We show that whilst existing tools can often effectively replicate known routes in familiar chemical spaces, they tend to fail in less explored regions. To overcome this challenge, we present a novel approach that integrates under-utilised reactions, resulting in improved performance in nearly 70% of tested cases. We introduce a tool that identifies, prioritises, and incorporates entirely new chemical transformations, demonstrating an average reduction in predicted synthesis costs of [~] 20%. Importantly, our tool can highlight new chemical reactions that should be prioritised for development to improve synthesis efficiency. Our findings lay the groundwork for next generation retrosynthesis tools, which will be capable of driving the discovery of novel chemical transformations and more efficient exploration of chemical space.

bioinformatics↗

Testing the Limits of SMILES-based De Novo Molecular Generation with Curriculum and Deep Reinforcement Learning

1Deep reinforcement learning methods have been shown to be potentially powerful tools for de novo design. Recurrent neural network (RNN)-based techniques are the most widely used methods in this space. In this work, we examine the behaviour of RNN-based methods when there are few (or no) examples of molecules with the desired properties in the training data. We find that targeted molecular generation is often possible, but the diversity of generated molecules is often reduced, and it is not possible to control the composition of generated molecular sets. To help overcome these issues, we propose a new curriculum learning-inspired, recurrent Iterative Optimisation Procedure that enables the optimisation of generated molecules for seen and unseen molecular profiles and allows the user to control whether a molecular profile is explored or exploited. Using our method, we generate specific and diverse sets of molecules with up to 18 times more scaffolds than standard methods for the same sample size. However, our results also point to significant limitations of one-dimensional molecular representations as used in this space. We find that the success or failure of a given molecular optimisation problem depends on the choice of SMILES.

bioinformatics↗