Search bioRxivSearch

Biology subjects

Tom Kodamullil, A.

Publications and source records attributed to Tom Kodamullil, A..

2 recordsLinked to original sources

STonKGs: A Sophisticated Transformer Trained on Biomedical Text and Knowledge Graphs

The majority of biomedical knowledge is stored in structured databases or as unstructured text in scientific publications. This vast amount of information has led to numerous machine learning-based biological applications using either text through natural language processing (NLP) or structured data through knowledge graph embedding models (KGEMs). However, representations based on a single modality are inherently limited. To generate better representations of biological knowledge, we propose STonKGs, a Sophisticated Transformer trained on biomedical text and Knowledge Graphs. This multimodal Transformer uses combined input sequences of structured information from KGs and unstructured text data from biomedical literature to learn joint representations. First, we pre-trained STonKGs on a knowledge base assembled by the Integrated Network and Dynamical Reasoning Assembler (INDRA) consisting of millions of text-triple pairs extracted from biomedical literature by multiple NLP systems. Then, we benchmarked STonKGs against two baseline models trained on either one of the modalities (i.e., text or KG) across eight different classification tasks, each corresponding to a different biological application. Our results demonstrate that STonKGs outperforms both baselines, especially on the more challenging tasks with respect to the number of classes, improving upon the F1-score of the best baseline by up to 0.083. Additionally, our pre-trained model as well as the model architecture can be adapted to various other transfer learning applications. Finally, the source code and pre-trained STonKGs models are available at https://github.com/stonkgs/stonkgs and https://huggingface.co/stonkgs/stonkgs-150k.

bioinformatics

The COVID-19 PHARMACOME: Rational Selection of Drug Repurposing Candidates from Multimodal Knowledge Harmonization

The SARS-CoV-2 pandemic has challenged researchers at a global scale. The scientific communitys massive response has resulted in a flood of experiments, analyses, hypotheses, and publications, especially in the field of drug repurposing. However, many of the proposed therapeutic compounds obtained from SARS-CoV-2 specific assays are not in agreement and thus demonstrate the need for a singular source of COVID-19 related information from which a rational selection of drug repurposing candidates can be made. In this paper, we present the COVID-19 PHARMACOME, a comprehensive drug-target-mechanism graph generated from a compilation of 10 separate disease maps and sources of experimental data focused on SARS-CoV-2 / COVID-19 pathophysiology. By applying our systematic approach, we were able to predict the synergistic effect of specific drug pairs, such as Remdesivir and Thioguanosine or Nelfinavir and Raloxifene, on SARS-CoV-2 infection. Experimental validation of our results demonstrate that our graph can be used to not only explore the involved mechanistic pathways, but also to identify novel combinations of drug repurposing candidates.

bioinformatics