Search bioRxiv⌕ Search

Biology subjects

Sobhan, M.

Publications and source records attributed to Sobhan, M..

4 recordsLinked to original sources

EMoMiS: A Pipeline for Epitope-based Molecular Mimicry Search in Protein Structures with Applications to SARS-CoV-2

MotivationEpitope-based molecular mimicry occurs when an antibody cross-reacts with two different antigens due to structural and chemical similarities. Molecular mimicry between proteins from two viruses can lead to beneficial cross-protection when the antibodies produced by exposure to one also react with the other. On the other hand, mimicry between a protein from a pathogen and a human protein can lead to auto-immune disorders if the antibodies resulting from exposure to the virus end up interacting with host proteins. While cross-protection can suggest the possible reuse of vaccines developed for other pathogens, cross-reaction with host proteins may explain side effects. There are no computational tools available to date for a large-scale search of antibody cross-reactivity. ResultsWe present a comprehensive Epitope-based Molecular Mimicry Search (EMoMiS) pipeline for computational molecular mimicry searches. EMoMiS, when applied to the SARS-CoV-2 Spike protein, identified eight examples of molecular mimicry with viral and human proteins. These findings provide possible explanations for (a) differential severity of COVID-19 caused by cross-protection due to prior vaccinations and/or exposure to other viruses, and (b) commonly seen COVID-19 side effects such as thrombocytopenia and thrombophilia. Our findings are supported by previously reported research but need validation with laboratory experiments. The developed pipeline is generic and can be applied to find mimicry for novel pathogens. It has applications in improving vaccine design. AvailabilityThe developed Epitope-based Molecular Mimicry Search Pipeline (EMoMiS) is available from https://biorg.cs.fiu.edu/emomis/. Contactgiri@cs.fiu.edu

bioinformatics↗

Quantifying Intratumor Heterogeneity by Key Genes Selected Using Concrete Autoencoder

The tumor cell population in cancer tissue has distinct molecular characteristics and exhibits different phenotypes, thus, resulting in different subpopulations. This phenomenon is known as Intratumor Heterogeneity (ITH), a major contributor to drug resistance, poor prognosis, etc. Therefore, quantifying the levels of ITH in cancer patients is essential, and many algorithms do so in different ways, using different types of omics data. DEPTH (Deviating gene Expression Profiling Tumor Heterogeneity) is the latest algorithm that uses transcriptomic data to evaluate the ITH score. It shows promising performance, has strong similarity with six other algorithms and has an advantage over two algorithms that uses the same type of data (tITH, sITH). However, it has a major drawback since it uses expression values of all the genes ([~]20K genes) in quantifying ITH levels. We hypothesize that a subset of key genes is sufficient to quantify the ITH level. To prove our hypothesis, we developed a deep learning-based computational framework using unsupervised Concrete Autoencoder (CAE) to select a set of cancer-specific key genes that can be used to evaluate the ITH score. For the experiment, we used gene expression profile data of tumor cohorts of breast, kidney, and lung cancer from the TCGA repository. Using multi-run CAE, we selected three sets of key genes, each set related to breast, kidney, and lung tumor cohorts. For the three cancers stated and three molecular subtypes of lung cancer, we calculated the ITH level using all genes and key genes selected by CAE and performed a side-by-side comparison. We could reach similar conclusions for survival and prognostic outcomes based on ITH scores derived from all genes and the sets of key genes. Additionally, for subtypes of lung cancer, the comparative distribution of ITH scores derived from all and key genes remains similar. Based on these observations, it can be stated that a subset of key genes, instead of all genes, is sufficient for ITH quantification. Our results also showed that many key genes are prognostically significant, which can be used as possible therapeutic targets.

bioinformatics↗

Molecular mimicry between Spike and human thrombopoietin may induce thrombocytopenia in COVID-19

SARS-CoV-2 causes COVID-19, a disease curiously resulting in varied symptoms and outcomes, ranging from asymptomatic to fatal. Autoimmunity due to cross-reacting antibodies resulting from molecular mimicry between viral antigens and host proteins may provide an explanation. We computationally investigated molecular mimicry between SARS-CoV-2 Spike and known epitopes. We discovered molecular mimicry hotspots in Spike and highlight two examples with tentative autoimmune potential and implications for understanding COVID-19 complications. We show that a TQLPP motif in Spike and thrombopoietin shares similar antibody binding properties. Antibodies cross-reacting with thrombopoietin may induce thrombocytopenia, a condition observed in COVID-19 patients. Another motif, ELDKY, is shared in multiple human proteins such as PRKG1 and tropomyosin. Antibodies cross-reacting with PRKG1 and tropomyosin may cause known COVID-19 complications such as blood-clotting disorders and cardiac disease, respectively. Our findings illuminate COVID-19 pathogenesis and highlight the importance of considering autoimmune potential when developing therapeutic interventions to reduce adverse reactions.

bioinformatics↗

Multi-run Concrete Autoencoder to Identify Prognostic lncRNAs for 12 Cancers

Long non-coding RNA plays a vital role in changing the expression profiles of various target genes that leads to cancer development. Thus, identifying prognostic lncRNAs related to different cancers might help in developing cancer therapy. To discover the critical lncRNAs that can identify the origin of different cancers, we proposed to use the state-of-the-art deep learning algorithm Concrete Autoencoder (CAE) in an unsupervised setting, which efficiently identifies a subset of the most informative features. However, CAE does not identify reproducible features in different runs due to its stochastic nature. We proposed a multi-run CAE (mrCAE) to identify a stable set of features to address this issue. The assumption is that a feature appearing in multiple runs carries more meaningful information about the data under consideration. The genome-wide lncRNA expression profiles of 12 different types of cancers, a total of 4,768 samples available in The Cancer Genome Atlas (TCGA), were analyzed to discover the key lncRNAs. The lncRNAs identified by multiple runs of CAE were added to the final list of key lncRNAs, which are capable of identifying 12 different cancers. Our results showed that mrCAE performs better in feature selection than single-run CAE, standard autoencoder (AE), and other state-of-the-art feature selection techniques. This study discovered a set of top-ranking 128 lncRNAs that could identify the origin of 12 different cancers with an accuracy of 95%. Survival analysis showed that 76 of 128 lncRNAs have the prognostic capability in differentiating high- and low-risk groups of patients in different cancers. The proposed mrCAE outperformed AE, which select latent features and were thought to be the best tools for dimensionality reduction. By selecting actual features instead of pseudo-features, mrCAE can be valuable for precision medicine. The identified prognostic lncRNAs (this work), mRNAs, miRNAs, and DNA methylated genes (future work) can lead to biomarkers and therapies for different cancers.

bioinformatics↗