Search bioRxiv⌕ Search

Biology subjects

Patil, R. S.

Publications and source records attributed to Patil, R. S..

2 recordsLinked to original sources

DeePNAP: A deep learning method to predict protein-nucleic acids binding affinity from sequence

Predicting the protein-nucleic acid (PNA) binding affinity solely from their sequences is of paramount importance for the experimental design and analysis of PNA interactions (PNAIs). A large number of currently developed models for binding affinity prediction are limited to specific PNAIs, while also relying on both sequence and structural information of the PNA complexes for both train/test and also as inputs. As PNA complex structures available are scarce, this significantly limits the diversity and generalizability due to a small training dataset. Additionally, a majority of the tools predict a single parameter such as binding affinity or free energy changes upon mutations, rendering a model less versatile for usage. Hence, we propose DeePNAP, a machine learning-based model trained on a vast and heterogeneous dataset with 14,401 entries (from both eukaryotes and prokaryotes) of ProNAB database, consisting of wild-type and mutant PNA complex binding parameters. Our model precisely predicts the binding affinity and free energy changes due to the mutation(s) of PNAIs exclusively from the sequences. While other similar tools extract features from both sequence and structure information, DeePNAP employs sequence-based features to yield high correlation coefficients between the predicted and experimental values with low root mean squared errors for PNA complexes in predicting the KD and {Delta}{Delta}G implying the generalizability of DeePNAP. Additionally, we have also developed a web interface hosting DeePNAP that can serve as a powerful tool to rapidly predict binding affinities for a myriad of PNAIs with high precision toward developing a deeper understanding of their implications in various biological systems. Web interface: http://14.139.174.41:8080/

biophysics↗

Molecular complex detection in protein interaction networks through reinforcement learning

Many, if not most, proteins assemble into higher-order complexes to perform their biological functions. Such protein-protein interactions (PPI) are often experimentally measured for pairs of proteins and summarized in a weighted PPI network, to which community detection algorithms can be applied to define the various higher-order protein complexes. Current methods, which include both unsupervised and supervised approaches, often assume that protein complexes manifest only as dense subgraphs, and in the case of supervised approaches, focus only on learning which subgraphs correspond to complexes, not how to find them in a network, a task that is currently solved using heuristics. However, learning to walk trajectories on a network with the goal of finding protein complexes lends itself naturally to a reinforcement learning (RL) approach, a strategy that has not been extensively explored for community detection. Here, we evaluated the use of a reinforcement learning pipeline for community detection in weighted protein-protein interaction networks to detect new protein complexes. Using known complexes, the algorithm is trained to calculate the value of different possible subgraph densities in the process of walking on the network to find a protein complex. Then, a distributed prediction algorithm scales the RL pipeline to search for protein complexes on large PPI networks. The reinforcement learning pipeline applied to a human PPI network consisting of 8k proteins and 60k PPI results in 1,157 protein complexes and shows competitive accuracy with improved speed when compared to previous algorithms. We highlight protein complexes harboring minimally characterized proteins including C4orf19, C18orf21, and KIAA1522, suggest TMC04 to be a putative additional subunit of the KICSTOR complex, and confirm the participation of C15orf41 in a higher-order complex with CDAN1, ASF1A, and HIRA by 3D structural modeling. Reinforcement learning offers several distinct advantages for community detection, including scalability and knowledge of the walk trajectories defining those communities.

bioinformatics↗