Search bioRxivSearch

Biology subjects

Gonzales-Diaz, H.

Publications and source records attributed to Gonzales-Diaz, H..

2 recordsLinked to original sources

Prediction of breast cancer proteins using molecular descriptors and artificial neural networks: a focus on cancer immunotherapy proteins, metastasis driver proteins, and RNA-binding proteins

BackgroundBreast cancer (BC) is a heterogeneous disease characterized by an intricate interplay between different biological aspects such as ethnicity, genomic alterations, gene expression deregulation, hormone disruption, signaling pathway alterations and environmental determinants. Due to the complexity of BC, the prediction of proteins involved in this disease is a trending topic in drug design. MethodsThis work is proposing accurate prediction classifier for BC proteins using six sets of protein sequence descriptors and 13 machine learning methods. After using a univariate feature selection for the mix of five descriptor families, the best classifier was obtained using multilayer perceptron method (artificial neural network) and 300 features. ResultsThe performance of the model is demonstrated by the area under the receiver operating characteristics (AUROC) of 0.980 {+/-} 0.0037 and accuracy of 0.936 {+/-} 0.0056 (3-fold cross-validation). Regarding the prediction of 4504 cancer-associated proteins using this model, the best ranked cancer immunotherapy proteins related to BC were RPS27, SUPT4H1, CLPSL2, POLR2K, RPL38, AKT3, CDK3, RPS20, RASL11A and UBTD1; the best ranked metastasis driver proteins related to BC were S100A9, DDA1, TXN, PRNP, RPS27, S100A14, S100A7, MAPK1, AGR3 and NDUFA13; and the best ranked RNA-binding proteins related to BC were S100A9, TXN, RPS27L, RPS27, RPS27A, RPL38, MRPL54, PPAN, RPS20 and CSRP1. ConclusionsThis powerful model predicts several BC-related proteins which should be deeply studied to find new biomarkers and better therapeutic targets. The script and the results are available as a free repository at https://github.com/muntisa/neural-networks-for-breast-cancer-proteins.

bioinformatics

Prediction of druggable proteins using machine learning and functional enrichment analysis: a focus on cancer-related proteins and RNA-binding proteins

BackgroundDruggable proteins are a trending topic in drug design. The druggable proteome can be defined as the percentage of proteins that have the capacity to bind an antibody or small molecule with adequate chemical properties and affinity. The screening and in silico modeling are critical activities for the reduction of experimental costs.\n\nMethodsThe current work proposes a unique prediction model for druggable proteins using amino acid composition descriptors of protein sequences and 13 machine learning linear and non-linear classifiers. After feature selection, the best classifier was obtained using the support vector machine method and 200 tri-amino acid composition descriptors.\n\nResultsThe high performance of the model is determined by an area under the receiver operating characteristics (AUROC) of 0.975 {+/-} 0.003 and accuracy of 0.929 {+/-} 0.006 (3-fold cross-validation). Regarding the prediction of cancer-associated proteins using this model, the best ranked druggable predicted proteins in the breast cancer protein set were CDK4, AP1S1, POLE, HMMR, RPL5, PALB2, TIMP1, RPL22, NFKB1 and TOP2A; in the cancer-driving protein set were TLL2, FAM47C, SAGE1, HTR1E, MACC1, ZFR2, VMA21, DUSP9, CTNNA3 and GABRG1; and in the RNA-binding protein set were PLA2G1B, CPEB2, NOL6, LRRC47, CTTN, CORO1A, SCAF11, KCTD12, DDX43 and TMPO.\n\nConclusionsThis powerful model predicts several druggable proteins which should be deeply studied to find better therapeutic targets and thus improve clinical trials. The scripts are freely available at https://github.com/muntisa/machine-learning-for-druggable-proteins.

bioinformatics