Search bioRxivSearch

Biology subjects

Jensen, L. J.

Publications and source records attributed to Jensen, L. J..

10 recordsLinked to original sources

CoCoScore: Context-aware co-occurrence scoring for text mining applications using distant supervision

Information extraction by mining the scientific literature is key to uncovering relations between biomedical entities. Most existing approaches based on natural language processing extract relations from single sentence-level co-mentions, ignoring co-occurrence statistics over the whole corpus. Existing approaches counting entity co-occurrences ignore the textual context of each co-occurrence. We propose a novel corpus-wide co-occurrence scoring approach to relation extraction that takes the textual context of each co-mention into account. Our method, called CoCoScore, scores the certainty of stating an association for each sentence that co-mentions two entities. CoCoScore is trained using distant supervision based on a gold-standard set of associations between entities of interest. Instead of requiring a manually annotated training corpus, co-mentions are labeled as positives/negatives according to their presence/absence in the gold standard. We show that CoCoScore outperforms previous approaches in identifying human disease-gene and tissue-gene associations as well as in identifying physical and functional protein-protein associations in different species. CoCoScore is a versatile text-mining tool to uncover pairwise associations via co-occurrence mining, within and beyond biomedical applications. CoCoScore is available at: https://github.com/JungeAlexander/cocoscore

bioinformatics

Cytoscape stringApp: Network analysis and visualization of proteomics data

Protein networks have become a popular tool for analyzing and visualizing the often long lists of proteins or genes obtained from proteomics and other high-throughput technologies. One of the most popular sources of such networks is the STRING database, which provides protein networks for more than 2000 organisms, including both physical interactions from experimental data and functional associations from curated pathways, automatic text mining, and prediction methods. However, its web interface is mainly intended for inspection of small networks and their underlying evidence. The Cytoscape software, on the other hand, is much better suited for working with large networks and offers greater flexibility in terms of network analysis, import and visualization of additional data. To include both resources in the same workflow, we created stringApp, a Cytoscape app that makes it easy to import STRING networks into Cytoscape, retains the appearance and many of the features of STRING, and integrates data from associated databases. Here, we introduce many of the stringApp features and show how they can be used to carry out complex network analysis and visualization tasks on a typical proteomics dataset, all through the Cytoscape user interface. stringApp is freely available from the Cytoscape app store: http://apps.cytoscape.org/apps/stringapp.

bioinformatics

Viruses.STRING: A virus-host protein-protein interaction database

As viruses continue to pose risks to global health, having a better un-derstanding of virus-host protein-protein interactions aids in the development of treatments and vaccines. Here, we introduce Viruses.STRING, a protein-protein interaction database specifically catering to virus-virus and virus-host interactions. This database combines evidence from experimental and text-mining channels to provide combined probabilities for interactions between viral and host proteins. The database contains 177,425 interactions between 239 viruses and 319 hosts. The database is publicly available at viruses.string-db.org, and the interaction data can also be accessed through the latest version of the Cytoscape STRING app.

bioinformatics

A Comprehensive Approach Characterizing Fusion Proteins and Their Interactions Using Biomedical Literature

Todays increase in scientific literature requires the efficient methods of data mining for improving the extraction of the useful information from texts. In this manuscript, we used a data and text mining method to identify fusions and their protein-protein interactions from published biomedical text. The extracted fusion proteins and their protein-protein interactions are used as a training set for a Naive Bayes classifier that is further used for final identification of testing dataset, consisting of 1817 fusions. Our method has a literature corpus, text and annotation mappers; keywords, rule bases, negative tokens, and pattern extractor; synonym tagger, normalization, regular expression mapper; and Naive Bayes classifier. We classified 1817 unique fusion proteins and their corresponding 2908 protein-protein interactions for 18 cancer types. Therefore, it can be used for screening literature for identifying mentions unique cases of fusions that can be further used for downstream analysis. It is available at http://protfus.md.biu.ac.il/.

bioinformatics

Text mining of disease-lifestyle associations to explain comorbidities in electronic health registries

Mining of electronic health registries can reveal vast numbers of disease correlations (from hereon referred to as comorbidities for simplicity). However, the underlying causes can be hard to identify, in part because health registries usually do not record important lifestyle factors such as diet, substance consumption, and physical activity. To address this challenge, I developed a text-mining approach that uses dictionaries of diseases and lifestyle factors for named entity recognition and subsequently for co-occurrence extraction of disease-lifestyle associations from Medline. I show that this approach is able to extract many correct associations and provide proof-of-concept that these can provide plausible explanations for comorbidities observed in Swedish and Danish health registry data.

bioinformatics

Text mining of 15 million full-text scientific articles

Across academia and industry, text mining has become a popular strategy for keeping up with the rapid growth of the scientific literature. Text mining of the scientific literature has mostly been carried out on collections of abstracts, due to their availability. Here we present an analysis of 15 million English scientific full-text articles published during the period 1823-2016. We describe the development in article length and publication sub-topics during these nearly 250 years. We showcase the potential of text mining by extracting published protein-protein, disease-gene, and protein subcellular associations using a named entity recognition system, and quantitatively report on their accuracy using gold standard benchmark data sets. We subsequently compare the findings to corresponding results obtained on 16.5 million abstracts included in MEDLINE and show that text mining of full-text articles consistently outperforms using abstracts only.

bioinformatics

An integrative method to unravel the host-parasite interactome: An orthology-based approach

The study of molecular host-parasite interactions is essential to understand parasitic infection and adaptation within the host system. As well, prevention and treatment of infectious diseases require clear understanding of the molecular crosstalk between parasites and their hosts. As yet, experimental large-scale identification of host-parasite molecular interactions remains challenging and the use of in silico predictions becomes then necessary. Here, we propose a computational integrative approach to predict host-parasite protein-protein interaction (PPI) networks resulting from the infection of human by 12 different parasites. We used an orthology-based method to transfer high-confidence intra-species interactions obtained from the STRING database to the corresponding inter-species protein pairs in the host-parasite system. To reduce the number of spurious predictions, our approach uses either the parasites predicted secretome and membrane proteins or only the secretome depending on whether they are uni- or multicellular respectively. Besides, the host proteome is filtered for proteins expressed in selected cellular localizations and tissues supporting the parasites growth. We evaluated the inferred interactions by analyzing the enriched biological processes and pathways in the predicted networks and their association to known parasitic invasion and evasion mechanisms. The resulting PPI networks were compared across parasites to identify common mechanisms that may define a global pathogenic hallmark. The predicted PPI networks can be visualized and downloaded at http://orthohpi.jensenlab.org.\n\nAuthor SummaryA protein-protein interaction (PPI) network is a collection of interactions between proteins from one or more organisms. Host-parasite PPIs are key to understanding the biology of different parasitic diseases, since predicting PPIs enable to know more about the parasite invasion, infection and persist. Our understanding of PPIs between host and parasites is still very limited, as no many systematic experimental studies have so far been performed. Efficacy of treatments for parasitic diseases is limited and in many cases parasites evolve resistance. Thus, there is an urgent need to develop novel drugs or vaccines for these neglected diseases, and thus interest in the functions and interactions of proteins associated with parasitism processes. Here we developed an in silico method to shed light on the interactome in twelve human parasites by combining an orthology based strategy and integrating, domain-domain interaction data, sub-cellular localization and to give a spatial context, we only took in to account those human tissues that support the parasites tropism. Here we show that is possible to identify relevant interactions across different parasites and their human host and that these interactions are well supported based on the biology of the parasites.

systems biology

Drug Target Ontology to Classify and Integrate Drug Discovery Data

BackgroundOne of the most successful approaches to develop new small molecule therapeutics has been to start from a validated druggable protein target. However, only a small subset of potentially druggable targets has attracted significant research and development resources. The Illuminating the Druggable Genome (IDG) project develops resources to catalyze the development of likely targetable, yet currently understudied prospective drug targets. A central component of the IDG program is a comprehensive knowledge resource of the druggable genome.\n\nResultsAs part of that effort, we have been developing a framework to integrate, navigate, and analyze drug discovery data based on formalized and standardized classifications and annotations of druggable protein targets, the Drug Target Ontology (DTO). DTO was constructed by extensive curation and consolidation of various resources. DTO classifies the four major drug target protein families, GPCRs, kinases, ion channels and nuclear receptors, based on phylogenecity, function, target development level, disease association, tissue expression, chemical ligand and substrate characteristics, and target-family specific characteristics. The formal ontology was built using a new software tool to auto-generate most axioms from a database while also supporting manual knowledge acquisition. A modular, hierarchical implementation facilitates development and maintenance and makes use of various external ontologies, thus integrating the DTO into the ecosystem of biomedical ontologies. As a formal OWL-DL ontology, DTO contains asserted and inferred axioms. Modeling data from the Library of Integrated Network-based Cellular Signatures (LINCS) program illustrates the potential of DTO for contextual data integration and nuanced definition of important drug target characteristics. DTO has been implemented in the IDG user interface Portal, Pharos and the TIN-X explorer of protein target disease relationships.\n\nConclusionsDTO was built based on the need for a formal semantic model for druggable targets including various related information such as protein, gene, protein domain, protein structure, binding site, small molecule drug, mechanism of action, protein tissue localization, disease association, and many other types of information. DTO will further facilitate the otherwise challenging integration and formal linking to biological assays, phenotypes, disease models, drug poly-pharmacology, binding kinetics and many other processes, functions and qualities that are at the core of drug discovery. The first version of DTO is publically available via the website http://drugtargetontology.org/, Github (https://github.com/DrugTargetOntology/DTO), and the NCBO Bioportal (https://bioportal.bioontology.org/ontologies/DTO). The long-term goal of DTO is to provide such an integrative framework and to populate the ontology with this information as a community resource.

bioinformatics

Tagger: BeCalm API for rapid named entity recognition

Most BioCreative tasks to date have focused on assessing the quality of text-mining annotations in terms of precision of recall. Interoperability, speed, and stability are, however, other important factors to consider for practical applications of text mining. The new BioCreative/BeCalm TIPS task focuses purely on these. To participate in this task, I implemented a BeCalm API within the real-time tagging server also used by the Reflect and EXTRACT tools. In addition to retrieval of patent abstracts, PubMed abstracts, and Pub-Med Central open-access articles as required in the TIPS task, the BeCalm API implementation facilitates retrieval of documents from other sources specified as custom request parameters. As in earlier tests, the tagger proved to be both highly efficient and stable, being able to consistently process requests of 5000 abstracts in less than half a minute including retrieval of the document text.

bioinformatics

EXTRACT 2.0: text-mining-assisted interactive annotation of biomedical named entities and ontology terms

1. INTRODUCTION 1. INTRODUCTION 2. EXPANDED SCOPE OF... FUNDING REFERENCES Databases increasingly rely on text-mining tools to support the curation process. The BioCreative interactive annotation task recently evaluated several such tools and found our tool EXTRACT to perform favorably in terms of usability and accelerated curation by 15-25% (Wang et al, 2016).\n\nThe original version of EXTRACT was designed to support annotation of metagenomic samples with semantically controlled environmental descriptors (Pafilis et al., 2016). For this reason, it focused on named entity recognition of terms from the Environment Ontology (ENVO) (Buttigieg et al., 2016) and ontologies relevant for describing host organisms (https://www.ncbi.nl ...

bioinformatics