Search bioRxiv⌕ Search

Biology subjects

Geffen, Y.

Publications and source records attributed to Geffen, Y..

2 recordsLinked to original sources

DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts

SummaryRecently, Deep Learning models, initially developed in the field of Natural Language Processing (NLP), were applied successfully to analyze protein sequences. A major drawback of these models is their size in terms of the number of parameters needed to be fitted and the amount of computational resources they require. Recently, "distilled" models using the concept of student and teacher networks have been widely used in NLP. Here, we adapted this concept to the problem of protein sequence analysis, by developing DistilProtBert, a distilled version of the successful ProtBert model. Implementing this approach, we reduced the size of the network and the running time by 50%, and the computational resources needed for pretraining by 98% relative to ProtBert model. Using two published tasks, we showed that the performance of the distilled model approaches that of the full model. We next tested the ability of DistilProtBert to distinguish between real and random protein sequences. The task is highly challenging if the composition is maintained on the level of singlet, doublet and triplet amino acids. Indeed, traditional machine learning algorithms have difficulties with this task. Here, we show that DistilProtBert preforms very well on singlet, doublet, and even triplet-shuffled versions of the human proteome, with AUC of 0.92, 0.91, and 0.87 respectively. Finally, we suggest that by examining the small number of false-positive classifications (i.e., shuffled sequences classified as proteins by DistilProtBert) we may be able to identify de-novo potential natural-like proteins based on random shuffling of amino acid sequences. Availabilityhttps://github.com/yarongef/DistilProtBert Contactyaron.geffen@biu.ac.il

bioinformatics↗

High-resolution lung adenocarcinoma expression subtypes identify tumors with dependencies on MET, CDK4, CDK6, and PD-L1

Lung adenocarcinoma is one of the most common cancer types with various treatment modalities. However, better biomarkers to predict therapeutic response are still needed to improve precision medicine. We utilized a consensus hierarchical clustering approach on 509 LUAD cases from TCGA to identify five robust LUAD expression subtypes. We then integrated genomic (patient and cell line) and proteomic data to help define biomarkers of response to targeted therapies and immunotherapies. This approach defined subtypes with unique proteogenomic and dependency profiles. S4-associated cell lines exhibited specific vulnerability to CDK6 and CDK6-cyclin D3 complex gene, CCND3. S3 was characterized by dependency on CDK4, immune-related expression patterns, and altered MET signaling; experimental validation showed that S3-associated cell lines responded to MET inhibitors, leading to increased PD-L1 expression. We further identified genomic features in S3 and S4 as biomarkers for enabling clinical diagnosis of these subtypes. Overall, our consensus hierarchical clustering approach identified robust tumor expression subtypes, and our subsequent integrative analysis of genomics, proteomics, and CRISPR screening data revealed subtype-specific biology and vulnerabilities. Our lung adenocarcinoma expression subtypes and their biomarkers could help identify patients likely to respond to CDK4/6, MET, or PD-L1 inhibitors, potentially improving patient outcome. SignificanceThrough integrative analysis of genomic, proteomic, and drug dependency data, we identified robust lung adenocarcinoma expression subtypes and found subtype-specific biomarkers of response, including CDK4/6, MET, and PD-L1 inhibitors.

cancer biology↗