bioRxiv · 10.1101/2023.12.10.570999
STRING-ing together protein complexes: extracting physical protein interactions from the literature
Abstract
Understanding biological processes relies heavily on curated knowledge of physical interactions between proteins. Yet, a notable gap remains between the information stored in databases of curated knowledge and the plethora of interactions documented in the scientific literature. To bridge this gap, we introduce ComplexTome, a manually annotated corpus designed to facilitate the development of text-mining methods for the extraction of complex formation relationships among biomedical entities. This corpus comprises 1,287 documents with [~]3, 500 relationships. We train a novel relation extraction model on this corpus and find that it can highly reliably identify physical protein interactions (F1-score=82.8%). We additionally enhance the models capabilities through unsupervised trigger word detection and apply it to extract relations and trigger words for these relations from all open publications in the domain literature. This information has been fully integrated into the latest version of the STRING database, and all introduced resources are openly accessible via Zenodo and GitHub.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mehryary, F., Nastou, K., Ohta, T., Jensen, L. J., Pyysalo, S.. 2023-12-11. STRING-ing together protein complexes: extracting physical protein interactions from the literature. https://doi.org/10.1101/2023.12.10.570999
Cite the original work for its findings. Save a collection to share your selection of sources.