Search bioRxivSearch

Biology subjects

Rasheed, K.

Publications and source records attributed to Rasheed, K..

2 recordsLinked to original sources

Deep evolutionary analysis reveals the design principles of fold A glycosyltransferases

Glycosyltransferases (GTs) are prevalent across the tree of life and regulate nearly all aspects of cellular functions by catalyzing synthesis of glycosidic linkages between diverse donor and acceptor substrates. Despite the availability of GT sequences from diverse organisms, the evolutionary basis for their complex and diverse modes of catalytic and regulatory functions remain enigmatic. Here, based on deep mining of over half a million GT-A fold sequences from diverse organisms, we define a minimal core component shared among functionally diverse enzymes. We find that variations in the common core and the emergence of hypervariable loops extending from the core contributed to the evolution of catalytic and functional diversity. We provide a phylogenetic framework relating diverse GT-A fold families for the first time and show that inverting and retaining mechanisms emerged multiple times independently during the course of evolution. We identify conserved modes of donor and acceptor recognition in evolutionarily divergent families and pinpoint the sequence and structural features for functional specialization. Using the evolutionary information encoded in primary sequences, we trained a machine learning classifier to predict donor specificity with nearly 88% accuracy and deployed it for the annotation of understudied GTs in five model organisms. Our studies provide an evolutionary framework for investigating the complex relationships connecting GT-A fold sequence, structure, function and regulation.

bioinformatics

Quantitative Structure-Mutation-Activity Relationship Tests (QSMART) Model for Protein Kinase Inhibitor Response Prediction

Predicting drug sensitivity profiles from genotypes is a major challenge in personalized medicine. Machine learning and deep neural network methods have shown promise in addressing this challenge, but the "black-box" nature of these methods precludes a mechanistic understanding of how and which genomic and proteomic features contribute to the observed drug sensitivity profiles. Here we provide a combination of statistical and neural network framework that not only estimates drug IC50 in cancer cell lines with high accuracy (R2 = 0.861 and RMSE = 0.818) but also identifies features contributing to the accuracy, thereby enhancing explainability. Our framework, termed QSMART, uses a multi-component approach that includes (1) collecting drug fingerprints, cancer cell lines multi-omics features, and drug responses, (2) testing the statistical significance of interaction terms, (3) selecting features by Lasso with Bayesian information criterion, and (4) using neural networks to predict drug response. We evaluate the contribution of each of these components and use a case study to explain the biological relevance of several selected features to protein kinase inhibitor response in non-small cell lung cancer cells. Specifically, we illustrate how interaction terms that capture associations between drugs and mutant kinases quantitatively contribute to the response of two EGFR inhibitors (afatinib and lapatinib) in non-small cell lung cancer cells. Although we have tested QSMART on protein kinase inhibitors, it can be extended across the proteome to investigate the complex relationships connecting genotypes and drug sensitivity profiles.

bioinformatics