Search bioRxivSearch

Biology subjects

Raghava, G. P. S.

Publications and source records attributed to Raghava, G. P. S..

6 recordsLinked to original sources

Prediction and analysis of skin cancer progression using genomics profiles of patients

Metastatic state of the Skin Cutaneous Melanoma (SKCM) has led to high mortality rate worldwide. Previously, various studies have revealed the association of the metastatic melanoma with the diminished survival rate in comparison to primary tumors. Thus, prediction of melanoma at primary tumor state is crucial to employ optimal therapeutic strategy for prolonged survival of patients. The RNA, miRNA and methylation data of The Cancer Genome Atlas (TCGA) cohort of SKCM is comprehensively analysed to recognize key genomic features that can categorize various states of metastatic tumors from primary tumors with high precision. Subsequently, various prediction models were developed using filtered genomic features implementing various machine learning techniques to classify these primary tumors from metastatic tumors. The SVC model (with class weight and RBF kernel) developed using 17 mRNA features achieved maximum MCC 0.73 with sensitivity, specificity and accuracy 89.19%, 90.48% and 89.47% respectively on independent validation dataset. Our study reveals that gene expression based features performs better than features obtained from miRNA profiling and epigenomic profiling. Our analysis shows that the expression of genes C7, MMP3, KRT14, KRT17, MASP1, and miRNA hsa-mir-205 and hsa-mir-203a are among the key genomic features that may substantially contribute to the oncogenesis of melanoma even on the basis of simple expression threshold. The major prediction models and analysis modules to predict metastatic and primary tumor samples of SKCM are available from a webserver, CancerSPP (http://webs.iiitd.edu.in/raghava/cancerspp/).

bioinformatics

Classification of early and late stage Liver Hepatocellular Carcinoma patients from their genomics and epigenomics profiles.

BackgroundLiver Hepatocellular Carcinoma (LIHC) is the second major cancer worldwide, responsible for millions of premature deaths every year. Prediction of clinical staging is vital to implement optimal therapeutic strategy and prognostic prediction in cancer patients. However, to date, no method has been developed for predicting stage of LIHC from genomic profile of samples.\n\nResultsIn current study, in silico models have been developed for classifying LIHC patients in early and late stage using RNA expression and DNA methylation data. The Cancer Genome Atlas (TCGA) dataset contains 173 early and 177 late stage samples of LIHC, was extensively analysed to identify differentially expressed RNA transcripts and methylated CpG sites that can discriminate early and late stages of LIHC samples with high precision. Naive Bayes model developed using 51 features that combine 21 CpG methylation sites and 30 RNA transcripts achieved maximum MCC 0.58 with accuracy 78.87% on validation dataset. Further, we also analysed genomics and epigenomics profiles of normal and LIHC samples and developed model to classify LIHC samples with AUROC 0.99. In addition, multiclass models developed for classifying samples in normal, early and late stage of cancer and achieved accuracy of 76.54% and AUROC of 0.86.\n\nConclusionOur study reveals stage prediction of LIHC samples with high accuracy based on genomics and epigenomics profiling is a challenging task in comparison to classification of LIHC and normal samples. Comprehensive analysis, differentially expressed RNA transcripts, methylated CpG sites in LIHC samples and prediction models are available from CancerLSP (http://webs.iiitd.edu.in/raghava/cancerlsp/).

bioinformatics

Computer-aided prediction of antigen presenting cell modulators for designing peptide-based vaccine adjuvants

BackgroundEvidences in literature strongly advocate the potential of immunomodulatory peptides for use as vaccine adjuvants. All the mechanisms of vaccine adjuvants ensuing immunostimulatory effects directly or indirectly stimulate Antigen Presenting Cells (APCs). While numerous methods have been developed in the past for predicting B-cell and T-cell epitopes; no method is available for predicting the peptides that can modulate the APCs.\n\nMethodsWe named the peptides that can activate APCs as A-cell epitopes and developed methods for their prediction in this study. A dataset of experimentally validated A-cell epitopes was collected and compiled from various resources. To predict A-cell epitopes, we developed Support Vector Machine-based machine learning models using different sequence-based features.\n\nResultsA hybrid model developed on a combination of sequence-based features (dipeptide composition and motif occurrence), achieved the highest accuracy of 96.91% with Matthews Correlation Coefficient (MCC) value of 0.94 on the training dataset. We also evaluated the hybrid models on an independent dataset and achieved a comparable accuracy of 94.93% with MCC 0.90.\n\nConclusionThe models developed in this study were implemented in a web-based platform VaxinPAD to predict and design immunomodulatory peptides or A-cell epitopes. This web server available at http://webs.iiitd.edu.in/raghava/vaxinpad/ and http://crdd.osdd.net/raghava/vaxinpad/ will facilitate researchers in designing peptide-based vaccine adjuvants.

bioinformatics

HumCFS: A database of fragile sites in human chromosomes

Genomic instability is the hallmark of cancer and several other pathologies, such as mental retardation; preferentially occur at specific loci in genome known as chromosomal fragile sites. HumCFS (http://webs.iiitd.edu.in/raghava/humcfs/) is a manually curated database provides comprehensive information on 118 experimentally characterized fragile sites present in human chromosomes. HumCFS comprises of 19068 entries with wide range of information such as nucleotide sequence of fragile sites, their length, coordinates on the chromosome, cytoband, their inducers and possibility of fragile site occurrence i.e. either rare or common etc. Each fragile region gene is further annotated to disease database DisGenNET, to understand its disease association. Protein coding genes are identified by annotating each fragile site to UCSC genome browser (GRCh38/hg38). To know the extent of miRNA lying in fragile site region, miRNA from miRBase has been mapped. Comprehensively, HumCFS encompasses mapping of 5010 genes with 19068 transcripts, 1104 miRNA and 3737 disease-associated genes on fragile sites. In order to facilitate users, we integrate standard web-based tools for easy data retrieval and analysis.

bioinformatics

Evaluation of protein-ligand docking methods on peptide-ligand complexes for docking small ligands to peptides

In the past, many benchmarking studies have been performed on protein-protein and protein-ligand docking however there is no study on peptide-ligand docking. In this study, we evaluated the performance of seven widely used docking methods (AutoDock, AutoDock Vina, DOCK 6, PLANTS, rDock, GEMDOCK and GOLD) on a dataset of 57 peptide-ligand complexes. Though these methods have been developed for docking ligands to proteins but we evaluate their ability to dock ligands to peptides. First, we compared TOP docking pose of these methods with original complex and achieved average RMSD from 4.74[A] for AutoDock to 12.63[A] for GEMDOCK. Next we evaluated BEST docking pose of these methods and achieved average RMSD from 3.82[A] for AutoDock to 10.83[A] for rDock. It has been observed that ranking of docking poses by these methods is not suitable for peptide-ligand docking as performance of their TOP pose is much inferior to their BEST pose. AutoDock clearly shows better performance compared to the other six docking methods based on their TOP docking poses. On the other hand, difference in performance of different docking methods (AutoDock, AutoDock Vina, PLANTS and DOCK 6) was marginal when evaluation was based on their BEST docking pose. Similar trend has been observed when performance is measured in terms of success rate at different cut-off values. In order to facilitate scientific community a web server PLDbench has been developed (http://webs.iiitd.edu.in/raghava/pldbench/).

bioinformatics

Prediction of residue-residue contacts in CASP12 targets from its predicted tertiary structures

One of the challenges in the field of structural proteomics is to predict residue-residue contacts in a protein. It is an integral part of CASP competitions due to its importance in the field of structural biology. This manuscript describes RRCPred 2.0 a method participated in CASP12 and predicted residue-residue contact in targets with high precision. In this approach, firstly 150 predicted protein structures were obtained from CASP12 Stage 2 tarball and ranked using clustering-based quality assessment software. Secondly, residue-residue contacts were assigned in top 10 protein structures based on distance between residues. Finally, residue-residue contacts were predicted in target protein based on consensus/average in top 10 predicted structures. This simple approach performs better than most of CASP12 methods in the categories of TBM and TBM/FM. It ranked 1st in following categories; i) TBM domain on list size L/5, ii) TBM/FM domain on list size L/5 and iii) TBM/FM domain on Top 10. These observations indicate that predicted tertiary structure of a protein can be used for predicting residue-residue contacts in protein with high accuracy.

bioinformatics