bioRxiv · 10.1101/2023.09.01.555843
Learning the sequence code for mRNA and protein abundance in human immune cells
Abstract
Accurate protein expression in human immune cells is essential for appropriate cellular function. The mechanisms that define protein abundance are complex and executed on transcriptional, post-transcriptional and post-translational level. Here, we present SONAR, a machine learning pipeline that learns the endogenous sequence code and that defines protein abundance in human cells. SONAR uses thousands of sequence features (SFs) to predict up to 63% of the protein abundance independently of promoter or enhancer information. SONAR uncovered the cell type-specific and activation-dependent usage of SFs. The deep knowledge of SONAR provides a map of biologically active SFs, which can be leveraged to manipulate the amplitude, timing, and cell type-specificity of protein expression. SONAR informed on the design of enhancer sequences to boost T cell receptor expression and to potentiate T cell function. Beyond providing fundamental insights in the regulation of protein expression, our study thus offers novel means to improve therapeutic and biotechnology applications. One Sentence SummarySONAR informs the design of cell type-specific protein expression in human cells
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Nicolet, B. P., Jurgens, A., Bresser, K., Guislain, A., Wolkers, M.. 2023-09-01. Learning the sequence code for mRNA and protein abundance in human immune cells. https://doi.org/10.1101/2023.09.01.555843
Cite the original work for its findings. Save a collection to share your selection of sources.