Search bioRxiv⌕ Search

Biology subjects

Kellman, B.

Publications and source records attributed to Kellman, B..

3 recordsLinked to original sources

Joint linear modeling of transcriptomics and proteomics is predictive of cancer metastasis

A central goal of conducting omics measurements is to understand how molecular features inform higher-order cell- and tissue-level phenotypes. In particular, multi-omics offers insights into how information encoded by the genome is coordinated through biological layers, resulting in functional outputs1. Due to myriad post-transcriptional regulatory processes, the coordination between mRNA and protein cannot be simply reduced to gene-wise correlation. Yet, both modalities have been shown to serve as representations of biological state, and multi-omics integration has been used to improve these representations. Multi-omics approaches typically do not focus on how mRNA and protein features coordinate, but rather use the additional information for improved prediction or feature selection. Here, instead, we showed that standard linear machine learning models provide an understanding of transcriptomic and proteomic coordination in the context of a biological phenotype of interest, in this case cancer metastasis. We find that, in the context of metastasis, a select subset of proteomic features--reflecting a more concentrated signal relative to the broadly distributed transcriptomic signal--offers additional information to that encoded by transcriptomics, as demonstrated by improved model performance when integrating the two modalities and the relative feature importance of proteomics. Top features show a depletion of gene-product overlap across modalities, indicating that the model primarily leverages instances in which the two modalities are providing complementary information with respect to phenotype. However, in instances when both modalities are selected for a given gene product, there is high information consistency that synergistically bolsters phenotype prediction. Altogether, by using model fits that relate both modalities to phenotype, we observe a nuanced coordination of protein and mRNA, in which both modalities tend to provide consistent information about phenotype, yet benefits remain to incorporating a combination of both complementary and reinforcing signals across modalities.

systems biology↗

Protein structure, a genetic encoding for glycosylation

Unlike DNA, RNA, and protein biosynthesis, dogma describes glycosylation as primarily determined by intrinsic cellular limitations, such as glycosyltransferase expression and precursor availability. However, this cannot explain the commonly-observed differences between glycans on the same protein. By examining site-specific glycosylation on diverse human proteins, we detected associations between protein structure and glycan structure, broadly generalizable to human-expressed glycoproteins. Through structural analysis of site-specific glycosylation data, we found protein-sequence and structural features consistently correlated with specific glycan features. To quantify these relationships, we present a new amino acid substitution matrix describing "glycoimpact", i.e., the association of primary protein structure and glycosylation. High-glycoimpact amino acids co-evolve with glycosites, and glycoimpact is high when estimates of amino acid conservation and variant pathogenicity diverge. We report thousands of disease variants near glycosites with high-glycoimpact, including several with known links to aberrant glycosylation (e.g., Oculocutaneous Albinism, Jakob-Creutzfeldt disease, Gerstmann-Straussler-Scheinker, and Gauchers Disease). Finally, glycoimpact quantification is validated by studying oligomannose-complex glycan ratios on HIV ENV, differential sialylation on IgG3 Fc, differential glycosylation on SARS-CoV-2 Spike, and fucose-modulated function of a tuberculosis monoclonal antibody. Finally, to test the causality of protein-glycan associations, we created 5 glycoimpact-designed novel Rituximab variants, 4 of which substantially changed glycoprofiles as predicted. In all, we report that site-specific glycan biosynthesis is influenced by underlying protein structure, enabling glycan structure prediction and genetic sequence-guided glycoengineering.

genetics↗

Decoding glycosylation potential from protein structure across human glycoproteins with a multi-view recurrent neural network

Glycosylation is described as a non-templated biosynthesis. Yet, the template-free premise is antithetical to the observation that different N-glycans are consistently placed at specific sites. It has been proposed that glycosite-proximal protein structures could constrain glycosylation and explain the observed microheterogeneity. Using site-specific glycosylation data, we trained a hybrid neural network to parse glycosites (recurrent neural network) and match them to feasible N-glycosylation events (graph neural network). From glycosite-flanking sequences, the algorithm predicts most human N-glycosylation events documented in the GlyConnect database and proposed structures corresponding to observed monosaccharide composition of the glycans at these sites. The algorithm also recapitulated glycosylation in Enhanced Aromatic Sequons, SARS-CoV-2 spike, and IgG3 variants, thus demonstrating the ability of the algorithm to predict both glycan structure and abundance. Thus, protein structure constrains glycosylation, and the neural network enables predictive in silico glycosylation of uncharacterized or novel protein sequences and genetic variants.

bioinformatics↗