Search bioRxiv⌕ Search

Biology subjects

Hoshen, Y.

Publications and source records attributed to Hoshen, Y..

2 recordsLinked to original sources

Detecting Anomalous Proteins Using Deep Representations

Many advances in biomedicine can be attributed to identifying unusual proteins and genes. Many of these proteins unique properties were discovered by manual inspection, which is becoming infeasible at the scale of modern protein datasets. Here, we propose to tackle this challenge using anomaly detection methods that automatically identify unexpected properties. We adopt a state-of-the-art anomaly detection paradigm from computer vision, to highlight unusual proteins. We generate meaningful representations without labeled inputs, using pretrained deep neural network models. We apply these protein language models (pLM) to detect anomalies in function, phylogenetic families, and segmentation tasks. We compute protein anomaly scores to highlight human prion-like proteins, distinguish viral proteins from their host proteome, and mark non-classical ion/metal binding proteins and enzymes. Other tasks concern segmentation of protein sequences into folded and unstructured regions. We provide candidates for rare functionality (e.g., prion proteins). Additionally, we show the anomaly score is useful in 3D folding-related segmentation. Our novel method shows improved performance over strong baselines and has objectively high performance across a variety of tasks. We conclude that the combination of pLM and anomaly detection techniques is a valid method for discovering a range of global and local protein characteristics.

bioinformatics↗

Biological representation disentanglement of single-cell data

Due to its internal state or external environment, a cells gene expression profile contains multiple signatures, simultaneously encoding information about its characteristics. Disentangling these factors of variations from single-cell data is needed to recover multiple layers of biological information and extract insight into the individual and collective behavior of cellular populations. While several recent methods were suggested for biological disentanglement, each has its limitations; they are either task-specific, cannot capture inherent nonlinear or interaction effects, cannot integrate layers of experimental data, or do not provide a general reconstruction procedure. We present biolord, a deep generative framework for disentangling known and unknown attributes in single-cell data. Biolord exposes the distinct effects of different biological processes or tissue structure on cellular gene expression. Based on that, biolord allows generating experimentally-inaccessible cell states by virtually shifting cells across time, space, and biological states. Specifically, we showcase accurate predictions of cellular responses to drug perturbations and generalization to predict responses to unseen drugs. Further, biolord disentangles spatial, temporal, and infection-related attributes and their associated gene expression signatures in a single-cell atlas of Plasmodium infection progression in the mouse liver. Biolord can handle partially labeled attributes by predicting a classification for missing labels, and hence can be used to computationally extend an infected hepatocyte population identified at a late stage of the infection to earlier stages. Biolord applies to diverse biological settings, is implemented using the scvi-tools library, and is released as open-source software at https://github.com/nitzanlab/biolord.

bioinformatics↗