Search bioRxiv⌕ Search

bioRxiv · 10.1101/2024.07.22.604688

Boosting the Predictive Power of Protein Representations with a Corpus of Text Annotations

Abstract

Protein language models are trained to predict amino acid sequences from vast protein databases, while learning to represent proteins as feature vectors. These vector representations have enabled impressive applications, from predicting mutation effects to protein folding. One of the reasons offered for the success of these models is that conserved sequence motifs tend to be important for protein fitness. Yet, the relationship between sequence conservation and fitness can be confounded by the evolutionary and environmental context. Should we therefore look to other data sources that may contain more direct functional information? In this work, we conduct a comprehensive study examining the effects of training protein models to predict nineteen types of text annotations from UniProt. Our results show that finetuning protein models on a subset of these annotations enhances the models predictive capabilities on a variety of function prediction tasks. Notably, our model outperforms the search algorithm BLAST, which none of the pre-trained protein models accomplished in our evaluation. Our results suggest that a much wider array of data modalities, such as text annotations, may be tapped to improve protein language models. We host our model checkpoints on https://huggingface.co/h4duan.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Duan, H., Skreta, M., Cotta, L., Rajaonson, E. M., Dhawan, N., Aspuru-Guzik, A., Maddison, C. J.. 2024-07-23. Boosting the Predictive Power of Protein Representations with a Corpus of Text Annotations. https://doi.org/10.1101/2024.07.22.604688

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

aaRSID, an engineered pyrrolysyl-tRNA synthetase platform for multi-probe proximity proteomics

Proximity labeling (PL) methods utilize spatially targeted chemical or enzymatic generation of a diffusible, reactive intermediate to covalently tag neighboring proteins in living systems. Unlike other tools for studying molecular interactions, PL can detect transient protein relationships with high spatial and temporal sensitivity, allowing for insight into their roles in biological processes. However, current enzymatic PL tools, such as TurboID and APEX2, are limited by their substrate structure and chemistry, which can generate significant background and/or perturb cellular physiology. To address these limitations, we have developed aminoacyl-tRNA synthetase ID (aaRSID), a PL tool that leverages an engineered pyrrolysyl tRNA synthetase (PylRS) for proximity labeling of proteins. We chose PylRS because it can catalyze promiscuous lysine labeling in the absence of its cognate tRNA and utilize a variety of non-canonical amino acids (ncAAs) as substrates. Here, we demonstrate aaRSID's intrinsic proximity labeling activity, use directed evolution to improve this activity, and apply the improved mutant (aaRSID-Ma1.3) for subcellular proteomics and multiplexed imaging. Our work establishes aminoacyl-tRNA synthetases as a new PL enzyme class and introduces a versatile chemical platform for developing ncAA-derived probes to map cellular microenvironments, greatly expanding the applications possible of PL technology.

biochemistry↗

Cellular uptake of folate-olaparib conjugates via folate receptor-mediated endocytosis: Potential for selective delivery of DNA damage response inhibitors into tumour cells

The folate receptor (FR) is overexpressed in a range of human tumours including ovarian cancer cells. We propose that the overexpression of the FR on the surface of ovarian tumour cells could be exploited for the selective delivery of a DNA damage response inhibitor (DDRi) in the form of an intact folate drug conjugate (FDC). This approach would improve the therapeutic index of the parent DDRi facilitating combination studies of the DDRi-based FDC with DNA damaging chemotherapy. FR-mediated cellular uptake of the proposed folate drug conjugates is requisite for FDC selective delivery into tumours. In this study, we synthesised a series of olaparib-based folate conjugates that maintained the biochemical PARP1 inhibition associated with olaparib and showed binding affinity for the folate receptor. Significantly, we identified compounds 10b and 11 that selectively enter FR overexpressing tumour cells via folate receptor-mediated endocytosis in their intact form and engage with their target as demonstrated by the potent inhibition of PARylation (KB cells, PARylation IC50 = 5.7 and 3.9 nM; respectively).

biochemistry↗

Architecture and Energy Transfer of the Bacterial Photosynthetic Unit

In phototrophic organisms, pigment-protein membrane complexes are densely packed to form photosynthetic units (PSUs) that capture solar energy and convert it into chemical energy. Although the structures of many individual photosynthetic complexes have been resolved, how they are arranged and interact with others within photosynthetic membranes to enable efficient excitation energy transfer (EET) remains poorly understood. Here, we report cryo-electron microscopy structures of PSU supercomplex assemblies from the phototrophic a-proteobacterium Rhodovulum viride, including an RC-LH1 core associated with one or two peripheral LH2 complexes and a curved LH2 tetramer. These membrane-derived assemblies define the relative positions and orientations of neighboring photosynthetic complexes and place their pigment arrays in proximity across antenna-antenna and antenna-core interfaces. Structure-based simulations identify potential EET pathways within the PSU assemblies and reveal rapid energy transfer across both LH2-LH2 and LH2-LH1 interfaces. Collectively, these findings provide insights into the assembly and structural modularity of bacterial PSUs and elucidate how the lateral organization of membrane protein complexes facilitates efficient energy transfer. This work extends structural studies of bacterial photosynthesis from individual complexes to their native higher-order assembly, providing a framework for understanding how photosynthetic supercomplex organization shapes energy migration and for guiding the design of artificial photosynthesis.

biochemistry↗