Search bioRxiv⌕ Search

Biology subjects

Dutton, O.

Publications and source records attributed to Dutton, O..

3 recordsLinked to original sources

Improving Inverse Folding models at Protein Stability Prediction without additional Training or Data

Deep learning protein sequence models have shown outstanding performance at de novo protein design and variant effect prediction. We substantially improve performance without further training or use of additional experimental data by introducing a second term derived from the models themselves which align outputs for the task of stability prediction. On a task to predict variants which increase protein stability the absolute success probabilities of PO_SCPLOWROTEINC_SCPLOWMPNN and ESMO_SCPLOWIFC_SCPLOW are improved by 11% and 5% respectively. We term these models PO_SCPLOWROTEINC_SCPLOWMPNN-O_SCPLOWDDC_SCPLOWG and ESMO_SCPLOWIFC_SCPLOW-O_SCPLOWDDC_SCPLOWG.

biophysics↗

NeRFax: An efficient and scalable conversion from the internal representation to Cartesian space

MotivationAccurate modelling of protein ensembles requires sampling of a large number of 3D conformations. A number of sampling approaches that use internal coordinates have been proposed, yet poor performance in the conversion from internal to Cartesian coordinates limits their applicability. ResultsWe describe here NeRFax, an efficient method for the conversion from internal to Cartesian coordinates that utilizes the platform-agnostic JAX Python library. The relative benefit of NeRFax is demonstrated here, on peptide chain reconstruction tasks. Our novel approach offers 35-175x times performance gains compared to previous state-of-the-art methods, whereas >10,000x speedup is reported in a reconstruction of a biomolecular condensate of 1,000 chains. AvailabilityNeRFax has purely open-source dependencies and is available at https://github.com/PeptoneLtd/nerfax. Contactoliver@peptone.io

bioinformatics↗

ADOPT: intrinsic protein disorder prediction through deep bidirectional transformers

Intrinsically disordered proteins (IDP) are important for a broad range of biological functions and are involved in many diseases. An understanding of intrinsic disorder is key to develop compounds that target IDPs. Experimental characterization of IDPs is hindered by the very fact that they are highly dynamic. Computational methods that predict disorder from the amino acid sequence have been proposed. Here, we present ADOPT, a new predictor of protein disorder. ADOPT is composed of a self-supervised encoder and a supervised disorder predictor. The former is based on a deep bidirectional transformer, which extracts dense residue level representations from Facebooks Evolutionary Scale Modeling (ESM) library. The latter uses a database of NMR chemical shifts, constructed to ensure balanced amounts of disordered and ordered residues, as a training and test dataset for protein disorder. ADOPT predicts whether a protein or a specific region is disordered with better performance than the best existing predictors and faster than most other proposed methods (a few seconds per sequence). We identify the features which are relevant for the prediction performance and show that good performance can already gained with less than 100 features. ADOPT is available as a standalone package at https://github.com/PeptoneLtd/ADOPT.

bioinformatics↗