Search bioRxivSearch

Biology subjects

Frank, A. T.

Publications and source records attributed to Frank, A. T..

3 recordsLinked to original sources

RNAPosers: Machine Learning Classifiers For RNA-Ligand Poses

Determining the 3-dimensional (3D) structures of ribonucleic acid (RNA)-small molecule complexes is critical to understanding molecular recognition in RNA. Computer docking can, in principle, be used to predict the 3D structure of RNA-small molecule complexes. Unfortunately, retrospective analysis has shown that the scoring functions that are typically used to rank poses tend to misclassify non-native poses as native, and vice versa. This misclassification of non-native poses severely limits the utility of computer docking in the context pose prediction, as well as in virtual screening. Here, we use machine learning to train a set of pose classifiers that estimate the relative \"nativeness\" of a set of RNA-ligand poses. At the heart of our approach is the use of a pose \"fingerprint\" that is a composite of a set of atomic fingerprints, which individually encode the local \"RNA environment\" around ligand atoms. We found that by ranking poses based on the classification scores from our machine learning classifiers, we were able to recover native-like poses better than when we ranked poses based on their docking scores. With a leave-one-out training and testing approach, we found that one of our classifiers could recover poses that were within 2.5 [A] of the native poses in [~]80% of the 88 cases we examined, and similarly, on a separate validation set, we could recover such poses in [~]70% of the cases. Our set of classifiers, which we refer to as RNAPosers, should find utility as a tool to aid in RNA-ligand pose prediction and so we make RNAPosers open to the academic community via https://github.com/atfrank/RNAPosers.

biophysics

Conditional Prediction of RNA Secondary Structure Using NMR Chemical Shifts

Inspired by methods that utilize chemical-mapping data to guide secondary structure prediction, we sought to develop a framework for using assigned chemical shift data to guide RNA secondary structure prediction. We first used machine learning to develop classifiers which predict the base-pairing status of individual residues in an RNA based on their assigned chemical shifts. Then, we used these base-pairing status predictions as restraints to guide RNA folding algorithms. Our results showed that we could recover the correct secondary folds for nearly all of the 108 RNAs in our dataset with remarkable accuracy. Finally, we assessed whether we could conditionally predict the structure of the model RNA, microRNA-20b (miR-20b), by folding it using folding restraints derived from chemical shifts associated with two distinct conformational states, one a free (apo) state and the other a protein-bound (holo) state. For this test, we found that by using folding restraints derived from chemical shifts, we could recover the two distinct structures of the miR-20b, confirming our ability to conditionally predict its secondary structure. A command-line tool for Chemical Shifts to Base-Pairing Status (CS2BPS) predictions in RNA has been incorporated into our CS2Structure Git repository and can be accessed via: https://github.com/atfrank/CS2Structure.

biophysics

Accelerating Dissociative Events in Molecular Dynamics Simulations by Selective Potential Scaling

Molecular dynamics (or MD) simulations can be a powerful tool for modeling complex dissociative processes such as ligand unbinding. However, many biologically relevant dissociative processes occur on timescales that far exceed the timescales of typical MD simulations. Here, we implement and apply an enhanced sampling method in which specific energy terms in the potential energy function are selectively "scaled" to accelerate dissociative events during simulations. Using ligand unbinding as an example of a complex dissociative process, we selectively scaled-up ligand-water interactions in an attempt to increase the rate of ligand unbinding. By applying our selectively scaled MD (or ssMD) approach to three cyclin-dependent kinase 2 (CDK2)-inhibitor complexes, we were able to significantly accelerate ligand unbinding thereby allowing, in some cases, unbinding events to occur within as little as 2 ns. Moreover, we found that we could make realistic estimates of the unbinding [Formula] as well as the binding free energies ({triangleup}Gsim) of the three inhibitors from our ssMD simulation data. To accomplish this, we employed a previously described Kramers-based rate extrapolation (KRE) method and a newly described free energy extrapolation (FEE) method. Because our ssMD approach is general, it should find utility as an easy-to-deploy, enhanced sampling method for modeling other dissociative processes.

biophysics