Search bioRxiv⌕ Search

Biology subjects

Nisonoff, H.

Publications and source records attributed to Nisonoff, H..

3 recordsLinked to original sources

Discovery and validation of the binding poses of allosteric fragment hits to PTP1b: From molecular dynamics simulations to X-ray crystallography

Fragment-based drug discovery has led to six approved drugs, but the small size of the chemical fragments used in such methods typically results in only weak interactions between the fragment and its target molecule, which makes it challenging to experimentally determine the three-dimensional poses fragments assume in the bound state. One computational approach that could help address this difficulty is long-timescale molecular dynamics (MD) simulation, which has been used in retrospective studies to recover experimentally known binding poses of fragments. Here, we present the results of long-timescale MD simulations that we used to prospectively discover binding poses for two series of fragments in allosteric pockets on a difficult and important pharmaceutical target, protein-tyrosine phosphatase 1b (PTP1b). Our simulations reversibly sampled the fragment association and dissociation process. One of the binding pockets found in the simulations has not to our knowledge been previously observed with a bound fragment, and the other pocket adopted a very rare conformation. We subsequently obtained high-resolution crystal structures of members of each fragment series bound to PTP1b, and the experimentally observed poses confirmed the simulation results. To the best of our knowledge, our findings provide the first demonstration that MD simulations can be used prospectively to determine fragment binding poses to previously unidentified pockets.

biophysics↗

Combining evolutionary and assay-labelled data for protein fitness prediction

Predictive modelling of protein properties has become increasingly important to the field of machine-learning guided protein engineering. In one of the two existing approaches, evolutionarily-related sequences to a query protein drive the modelling process, without any property measurements from the laboratory. In the other, a set of protein variants of interest are assayed, and then a supervised regression model is estimated with the assay-labelled data. Although a handful of recent methods have shown promise in combining the evolutionary and supervised approaches, this hybrid problem has not been examined in depth, leaving it unclear how practitioners should proceed, and how method developers should build on existing work. Herein, we present a systematic assessment of methods for protein fitness prediction when evolutionary and assay-labelled data are available. We find that a simple baseline approach we introduce is competitive with and often outperforms more sophisticated methods. Moreover, our simple baseline is plug-and-play with a wide variety of established methods, and does not add any substantial computational burden. Our analysis highlights the importance of systematic evaluations and sufficient baselines.

synthetic biology↗

Sparse Epistatic Regularization of Deep Neural Networks for Inferring Fitness Functions

Despite recent advances in high-throughput combinatorial mutagenesis assays, the number of labeled sequences available to predict molecular functions has remained small for the vastness of the sequence space combined with the ruggedness of many fitness functions. Expressive models in machine learning (ML), such as deep neural networks (DNNs), can model the nonlinearities in rugged fitness functions, which manifest as high-order epistatic interactions among the mutational sites. However, in the absence of an inductive bias, DNNs overfit to the small number of labeled sequences available for training. Herein, we exploit the recent biological evidence that epistatic interactions in many fitness functions are sparse; this knowledge can be used as an inductive bias to regularize DNNs. We have developed a method for sparse epistatic regularization of DNNs, called the epistatic net (EN), which constrains the number of non-zero coefficients in the spectral representation of DNNs. For larger sequences, where finding the spectral transform becomes computationally intractable, we have developed a scalable extension of EN, which subsamples the combinatorial sequence space uniformly inducing a sparse-graph-code structure, and regularizes DNNs using the resulting greedy optimization method. Results on several biological landscapes, from bacterial to protein fitness functions, show that EN consistently improves the prediction accuracy of DNNs and enables them to outperform competing models which assume other forms of inductive biases. EN estimates all the higher-order epistatic interactions of DNNs trained on massive sequence spaces--a computational problem that takes years to solve without leveraging the epistatic sparsity in the fitness functions.

bioinformatics↗