Search bioRxivSearch

Biology subjects

Kaur, H.

Publications and source records attributed to Kaur, H..

4 recordsLinked to original sources

Expression based biomarkers and models to classify early and late stage samples of Papillary Thyroid Carcinoma

In this study, we describe the key transcripts and machine learning models developed for classifying the early and late stage samples of Papillary Thyroid Cancer (PTC), using transcripts expression data from The Cancer Genome Atlas (TCGA). First, we rank all the transcripts on the basis of area under receiver operating characteristic curve, (AUROC) value to discriminate the early and late stage, based on an expression threshold. With the expression of a single transcript DCN, we can classify the stage samples with a 68.5% accuracy and AUROC of 0.66. Then we implemented various combination of multiple gene panels, selected using various gold standard feature selection techniques. The model based on the expression of 36 multiple transcripts (protein coding and non-coding) selected using SVC-L1 achieves the maximum accuracy of 74.51% with AUROC of 0.75 on independent validation dataset with balanced sensitivity and specificity. Further, these signatures also performed well on external microarray data obtained from GEO, predicting nearly 70% (12 samples out of 17 samples) early stage samples correctly. Further, multiclass model, classifying the normal, early and late stage samples achieves the accuracy of 75.43% with AUROC of 0.80 on independent validation dataset. With correlation analysis, we found that transcripts with maximum change in correlation of their expression in both the stages are significantly enriched in neuroactive ligand receptor interaction pathway. We also propose a panel of five protein coding transcripts, which on the basis of their expression, can segregate cancer and normal samples with 97.32% accuracy and AUROC of 0.99 on independent validation dataset. All the models and dataset used in this study are available from the web server CancerTSP (http://webs.iiitd.edu.in/raghava/cancertsp/).

bioinformatics

Prediction and analysis of skin cancer progression using genomics profiles of patients

Metastatic state of the Skin Cutaneous Melanoma (SKCM) has led to high mortality rate worldwide. Previously, various studies have revealed the association of the metastatic melanoma with the diminished survival rate in comparison to primary tumors. Thus, prediction of melanoma at primary tumor state is crucial to employ optimal therapeutic strategy for prolonged survival of patients. The RNA, miRNA and methylation data of The Cancer Genome Atlas (TCGA) cohort of SKCM is comprehensively analysed to recognize key genomic features that can categorize various states of metastatic tumors from primary tumors with high precision. Subsequently, various prediction models were developed using filtered genomic features implementing various machine learning techniques to classify these primary tumors from metastatic tumors. The SVC model (with class weight and RBF kernel) developed using 17 mRNA features achieved maximum MCC 0.73 with sensitivity, specificity and accuracy 89.19%, 90.48% and 89.47% respectively on independent validation dataset. Our study reveals that gene expression based features performs better than features obtained from miRNA profiling and epigenomic profiling. Our analysis shows that the expression of genes C7, MMP3, KRT14, KRT17, MASP1, and miRNA hsa-mir-205 and hsa-mir-203a are among the key genomic features that may substantially contribute to the oncogenesis of melanoma even on the basis of simple expression threshold. The major prediction models and analysis modules to predict metastatic and primary tumor samples of SKCM are available from a webserver, CancerSPP (http://webs.iiitd.edu.in/raghava/cancerspp/).

bioinformatics

Structure-activity relationship of flavin analogs that target the FMN riboswitch

The flavin mononucleotide (FMN) riboswitch is an emerging target for the development of novel RNA-targeting antibiotics. We previously discovered an FMN derivative --5FDQD-- that protects mice against diarrhea-causing Clostridium difficile bacteria. Here, we present the structure-based drug design strategy that led to the discovery of this fluoro-phenyl derivative with antibacterial properties. This approach involved the following stages: (1) structural analysis of all available free and bound FMN riboswitch structures; (2) design, synthesis and purification of derivatives; (3) in vitro testing for productive binding using two chemical probing methods; (4) in vitro transcription termination assays; (5) resolution of the crystal structures of the FMN riboswitch in complex with the most mature candidates. In the process, we delineated principles for productive binding to this riboswitch, thereby demonstrating the effectiveness of a coordinated structure-guided approach to designing drugs against RNA.\n\nGRAPHICAL ABSTRACT\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=146 SRC=\"FIGDIR/small/389148_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (41K):\norg.highwire.dtl.DTLVardef@f59798org.highwire.dtl.DTLVardef@1b38b02org.highwire.dtl.DTLVardef@6b50faorg.highwire.dtl.DTLVardef@1915fc4_HPS_FORMAT_FIGEXP M_FIG Exploring the chemical structure landscape of FMN riboswitch binders.\n\nC_FIG

biochemistry

Classification of early and late stage Liver Hepatocellular Carcinoma patients from their genomics and epigenomics profiles.

BackgroundLiver Hepatocellular Carcinoma (LIHC) is the second major cancer worldwide, responsible for millions of premature deaths every year. Prediction of clinical staging is vital to implement optimal therapeutic strategy and prognostic prediction in cancer patients. However, to date, no method has been developed for predicting stage of LIHC from genomic profile of samples.\n\nResultsIn current study, in silico models have been developed for classifying LIHC patients in early and late stage using RNA expression and DNA methylation data. The Cancer Genome Atlas (TCGA) dataset contains 173 early and 177 late stage samples of LIHC, was extensively analysed to identify differentially expressed RNA transcripts and methylated CpG sites that can discriminate early and late stages of LIHC samples with high precision. Naive Bayes model developed using 51 features that combine 21 CpG methylation sites and 30 RNA transcripts achieved maximum MCC 0.58 with accuracy 78.87% on validation dataset. Further, we also analysed genomics and epigenomics profiles of normal and LIHC samples and developed model to classify LIHC samples with AUROC 0.99. In addition, multiclass models developed for classifying samples in normal, early and late stage of cancer and achieved accuracy of 76.54% and AUROC of 0.86.\n\nConclusionOur study reveals stage prediction of LIHC samples with high accuracy based on genomics and epigenomics profiling is a challenging task in comparison to classification of LIHC and normal samples. Comprehensive analysis, differentially expressed RNA transcripts, methylated CpG sites in LIHC samples and prediction models are available from CancerLSP (http://webs.iiitd.edu.in/raghava/cancerlsp/).

bioinformatics