Search bioRxiv⌕ Search

Biology subjects

Devi, N. L.

Publications and source records attributed to Devi, N. L..

4 recordsLinked to original sources

A random forest model for predicting exosomal proteins using evolutionary information and motifs

Identification of secretory proteins in body fluids is one of the key challenges in the development of non-invasive diagnostics. It has been shown in the part that a significant number of proteins are secreted by cells via exosomes called exosomal proteins. In this study, an attempt has been made to build a model that can predict exosomal proteins with high precision. All models are trained, tested, and evaluated on a non-redundant dataset comprising 2831 exosomal and 2831 non-exosomal proteins, where no two proteins have more than 40% similarity. Initially, the standard similarity-based method BLAST was used to predict exosomal proteins, which failed due to low-level similarity in the dataset. To overcome this challenge, machine learning based models have been developed using compositional features of proteins and achieved highest AUROC of 0.70. The performance of the ML-based models improved significantly to AUROC of 0.73 when evolutionary information in the form of PSSM profiles was used for building models. Our analysis indicates that exosomal proteins have wide range of motifs. In addition, it was observed that exosomal proteins contain different types of sequence-based motifs, which can be used for predicting exosomal proteins. Finally, a hybrid method has been developed that combines a motif-based approach and an ML-based model for predicting exosomal proteins, achieving a maximum AUROC 0.85 and MCC of 0.56 on an independent dataset. The hybrid model in this study performs better than the presently available methods when assessed on an independent dataset. A web server and a standalone software ExoProPred has been created for the scientific community to provide service, code, and data. (https://webs.iiitd.edu.in/raghava/exopropred/). KeypointsO_LIExosomal proteins or non-classical secretory proteins are secreted by via exosomes C_LIO_LIA method has been developed for predicting exosomal proteins C_LIO_LIModels have been trained, tested, and evaluated on non-redundant dataset C_LIO_LIWide range of sequence motifs have been discovered in exosomal proteins C_LIO_LIA web server and standalone software have been developed C_LI

bioinformatics↗

A method for predicting linear and conformational B-cell epitopes in an antigen from its primary sequence

B-cell is an essential component of the immune system that plays a vital role in providing the immune response against any pathogenic infection by producing antibodies. Existing methods either predict linear or conformational B-cell epitopes in an antigen. In this study, a single method was developed for predicting both types (linear/conformational) of B-cell epitopes. The dataset used in this study contains 3875 B-cell epitopes and 3996 non-B-cell epitopes, where B-cell epitopes consist of both linear and conformational B-cell epitopes. Our primary analysis indicates that certain residues (like Asp, Glu, Lys, Asn) are more prominent in B-cell epitopes. We developed machine-learning based methods using different types of sequence composition and achieved the highest AUC of 0.80 using dipeptide composition. In addition, models were developed on selected features, but no further improvement was observed. Our similarity-based method implemented using BLAST shows a high probability of correct prediction with poor sensitivity. Finally, we came up with a hybrid model that combine alignment free (dipeptide based random forest model) and alignment-based (BLAST based similarity) model. Our hybrid model attained maximum AUC 0.83 with MCC 0.49 on the independent dataset. Our hybrid model performs better than existing methods on an independent dataset used in this study. All models trained and tested on 80% data using cross-validation technique and final model was evaluated on 20% data called independent or validation dataset. A webserver and standalone package named "CLBTope" has been developed for predicting, designing, and scanning B-cell epitopes in an antigen sequence (https://webs.iiitd.edu.in/raghava/clbtope/).

bioinformatics↗

Prediction and scanning of IL-5 inducing peptides using alignment-free and alignment-based method

Interleukin-5 (IL-5) is the key cytokine produced by T-helper, eosinophils, mast and basophils cells. It can act as an enticing therapeutic target due to its pivotal role in several eosinophil-mediated diseases. Though numerous methods have been developed to predict HLA binders and cytokines-inducing peptides, no method was developed for predicting IL-5 inducing peptides. All models in this study have been trained, tested and validated on experimentally validated 1907 IL-5 inducing and 7759 non-IL-5 inducing peptides obtained from IEDB. First, alignment-based methods have been developed using similarity and motif search. These alignment-based methods provide high precision but poor coverage. In order to overcome this limitation, we developed machine learning-based models for predicting IL-5 inducing peptides using a wide range of peptide features. Our random-forest model developed using selected 250 dipeptides achieved the highest performance among alignment-free methods with AUC 0.75 and MCC 0.29 on validation dataset. In order to improve the performance, we developed an ensemble or hybrid method that combined alignment-based and alignment-free methods. Our hybrid method achieved AUC 0.94 with MCC 0.60 on validation/ independent dataset. The best model developed in this study has been incorporated in the web server IL5pred (https://webs.iiitd.edu.in/raghava/il5pred/). Key PointsO_LIIL-5 is a regulatory cytokine that plays a vital role in eosinophil-mediated diseases C_LIO_LIBLAST-based similarity search against IL-5 inducing peptides was employed C_LIO_LIA hybrid approach combines alignment-based and alignment-free methods C_LIO_LIAlignment-free models are based on machine learning techniques C_LIO_LIA web server IL5pred and its standalone software have been developed C_LI Authors BiographyO_LIDr. Naorem Leimarembi Devi is currently working as a DBT-Research Associate in Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LINeelam Sharma is pursuing her Ph.D. in Computational Biology from the Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LIO_LIProf. G.P.S. Raghava is currently working as Professor and Head of Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India. C_LI

bioinformatics↗

Transcriptomics based prediction of metastasis in TNBC patients: Challenges in cross-platforms validation

Triple-negative breast cancer (TNBC) is more prone to metastasis and recurrence than other breast cancer subtypes. This study aimed to identify genes that can act as diagnostic biomarkers for predicting lymph node metastasis in TNBC patients. The transcriptomic data of TNBC with or without lymph node metastasis was acquired from TCGA, and the differentially expressed genes were identified. Further, logistic-regression method has been used to identify the top 15 genes (or 15 gene signatures) based on their ability to predict metastasis (AUC>0.65). These 15 gene signatures were used to develop machine learning techniques based prediction models; Gaussian Naive Bayes classifier outperformed other with AUC>0.80 on both training and validation datasets. The best model failed drastically on nine independent microarray datasets obtained from GEO. We investigated the reason for the failure of our best model, and it was observed that the certain genes in 15 gene signatures were showing opposite regulating trends, i.e., genes are upregulated in TCGA-TNBC patients while it is downregulated on other microarray datasets or vice-versa. In conclusion, the 15 gene signatures may act as diagnostic markers for the detection of lymph node metastatic status in TCGA dataset, but quite challenging across multiple platforms. We also identified the prognostic potential of the 15 selected genes and found that overexpression of ZNRF2, FRZB, and TCEAL4 was associated with poor survival with HR>2.3 and p-value[≤]0.05. In order to provide services to the scientific community, we developed a webserver named "MTNBCPred" for the prediction of metastatic and non-metastatic lymph node status of TNBC patients (http://webs.iiitd.edu.in/raghava/mtnbcpred/).

bioinformatics↗