bioRxiv · 10.1101/2020.08.04.236729
Genome-wide Prediction of Small Molecule Binding to Remote Orphan Proteins Using Distilled Sequence Alignment Embedding
Abstract
Endogenous or surrogate ligands of a vast number of proteins remain unknown. Identification of small molecules that bind to these orphan proteins will not only shed new light into their biological functions but also provide new opportunities for drug discovery. Deep learning plays an increasing role in the prediction of chemical-protein interactions, but it faces several challenges in protein deorphanization. Bioassay data are highly biased to certain proteins, making it difficult to train a generalizable machine learning model for the proteins that are dissimilar from the ones in the training data set. Pre-training offers a general solution to improving the model generalization, but needs incorporation of domain knowledge and customization of task-specific supervised learning. To address these challenges, we develop a novel protein pre-training method, DIstilled Sequence Alignment Embedding (DISAE), and a module-based fine-tuning strategy for the protein deorphanization. In the benchmark studies, DISAE significantly improves the generalizability and outperforms the state-of-the-art methods with a large margin. The interpretability analysis of pre-trained model suggests that it learns biologically meaningful information. We further use DISAE to assign ligands to 649 human orphan G-Protein Coupled Receptors (GPCRs) and to cluster the human GPCRome by integrating their phylogenetic and ligand relationships. The promising results of DISAE open an avenue for exploring the chemical landscape of entire sequenced genomes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cai, T., Lim, H., Abbu, K. A., Qiu, Y., Nussinov, R., Xie, L.. 2020-08-05. Genome-wide Prediction of Small Molecule Binding to Remote Orphan Proteins Using Distilled Sequence Alignment Embedding. https://doi.org/10.1101/2020.08.04.236729
Cite the original work for its findings. Save a collection to share your selection of sources.