Search bioRxiv⌕ Search

Biology subjects

Kogler-Anele, L.

Publications and source records attributed to Kogler-Anele, L..

2 recordsLinked to original sources

Active Learning for Protein Structure Prediction

Accurate protein structure prediction is challenging especially for families and types that are not well-represented in current training data. Active learning selects candidates for labeling with the aim of most rapidly improving model performance. In a general, many labeling strategies have been proposed. However, with protein structure prediction, most of these strategies dont apply or are difficult due to the high-dimensional, variable-dimensional regression target and the inherent complexity of the models involved. We applied a novel active learning strategy, DEW-DROP, to protein structure prediction on two different protein datasets: VHH-only antibodies (Nanobodies), and Mycobacterial proteins. We introduce a domain-specific fine-tuned Equifold model for VHH structures and apply DEWDROP to generate ensembles of predictions using Monte Carlo dropout. Using the statistics of these we select batches with high information content for labeling. We show that DEWDROP (1) improves model training efficiency through batch optimization outper-forming baselines, and (2) selects data with relevant high information content.

molecular biology↗

Deep Batch Active Learning for Drug Discovery

A key challenge in drug discovery is to optimize, in silico, various absorption and affinity properties of small molecules. One strategy that was proposed for such optimization process is active learning. In active learning molecules are selected for testing based on their likelihood of improving model performance. To enable the use of active learning with advanced neural network models we developed two novel active learning batch selection methods. These methods were tested on several public datasets for different optimization goals and with different sizes. We have also curated new affinity datasets that provide chronological information on state-of-the-art experimental strategy. As we show, for all datasets the new active learning methods greatly improved on existing and current batch selection methods leading to significant potential saving in the number of experiments needed to reach the same model performance. Our methods are general and can be used with any package including the popular DeepChem library.

bioinformatics↗