Search bioRxiv⌕ Search

Biology subjects

Colwell, L.

Publications and source records attributed to Colwell, L..

3 recordsLinked to original sources

ProteInfer: deep networks for protein functional inference

Predicting the function of a protein from its amino acid sequence is a long-standing challenge in bioinformatics. Traditional approaches use sequence alignment to compare a query sequence either to thousands of models of protein families or to large databases of individual protein sequences. Here we instead employ deep convolutional neural networks to directly predict a variety of protein functions - EC numbers and GO terms - directly from an unaligned amino acid sequence. This approach provides precise predictions which complement alignment-based methods, and the computational efficiency of a single neural network permits novel and lightweight software interfaces, which we demonstrate with an in-browser graphical interface for protein function prediction in which all computation is performed on the users personal computer with no data uploaded to remote servers. Moreover, these models place full-length amino acid sequences into a generalised functional space, facilitating downstream analysis and interpretation. To read the interactive version of this paper, please visit https://google-research.github.io/proteinfer/ O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=200 SRC="FIGDIR/small/461077v2_ufig1.gif" ALT="Figure 1"> View larger version (32K): org.highwire.dtl.DTLVardef@74f8f1org.highwire.dtl.DTLVardef@183ae0borg.highwire.dtl.DTLVardef@1754f3org.highwire.dtl.DTLVardef@1ca54f5_HPS_FORMAT_FIGEXP M_FIG C_FIG QR code for the interactive version of this preprint at https://google-research.github.io/proteinfer/

bioinformatics↗

Machine Learning Optimization of Photosynthetic Microbe Cultivation and Recombinant Protein Production

BackgroundArthrospira platensis (commonly known as spirulina) is a promising new platform for low-cost manufacturing of biopharmaceuticals. However, full realization of the platforms potential will depend on achieving both high growth rates of spirulina and high expression of therapeutic proteins. ObjectiveWe aimed to optimize culture conditions for the spirulina-based production of therapeutic proteins. MethodsWe used a machine learning approach called Bayesian black-box optimization to iteratively guide experiments in 96 photobioreactors that explored the relationship between production outcomes and 17 environmental variables such as pH, temperature, and light intensity. ResultsOver 16 rounds of experiments, we identified key variable adjustments that approximately doubled spirulina-based production of heterologous proteins, improving volumetric productivity between 70% to 100% in multiple bioreactor setting configurations. ConclusionAn adaptive, machine learning-based approach to optimize heterologous protein production can improve outcomes based on complex, multivariate experiments, identifying beneficial variable combinations and adjustments that might not otherwise be discoverable within high-dimensional data.

bioengineering↗

Using Single Protein/Ligand Binding Models to Predict Active Ligands for Unseen Proteins

Machine learning models that predict which small molecule ligands bind a single protein target report high levels of accuracy for held-out test data. An important challenge is to extrapolate and make accurate predictions for new protein targets. Improvements in drug-target interaction (DTI) models that address this challenge would have significant impact on drug discovery by eliminating the need for high-throughput screening experiments against new protein targets. Here we propose a data augmentation strategy that addresses this challenge to enable accurate prediction in cases where no experimental data is available. To proceed, we first build single protein-ligand binding models and use these models to predict whether additional ligands bind to each protein. We then use these predictions to augment the experimental data, train standard DTI models, and predict interactions between unseen test proteins and ligands. This approach achieves Area Under the Receiver Operator Characteristic (AUC) > 0.9 consistently on test sets consisting exclusively of proteins and ligands for which the model is given no experimental data. We verify that performance improvements extend to held-out test proteins distant from the training set. Our data augmentation framework can be applied to any DTI model, and enhances performance on a range of simple models.

bioinformatics↗