Search bioRxiv⌕ Search

Biology subjects

Demaray, J.

Publications and source records attributed to Demaray, J..

2 recordsLinked to original sources

Interpretable biophysical neural networks of transcriptional activation domains separate roles of protein abundance and coactivator binding

Deep neural networks have improved the accuracy of many difficult prediction tasks in biology, but it remains challenging to interpret these networks and learn molecular mechanisms. Here, we address the interpretability challenges associated with predicting transcriptional activation domains from protein sequence. Activation domains, regions within transcription factors that drive gene expression, were traditionally difficult to predict due to their sequence diversity and poor conservation. Multiple deep neural networks can now accurately predict activation domains, but these predictors are difficult to interpret. With the goal of interpretability, we designed simple neural networks that incorporated biophysical models of activation domains. The simplicity of these neural networks allowed us to visualize their parameters and directly interpret what the networks learned. The biophysical neural networks revealed two new ways that arrangement (i.e. the sequence grammar) of activation domain controlled function: 1) hydrophobic residues both increase activation domain strength and decrease protein abundance, and 2) acidic residues control both activation domain strength and protein abundance. Notably, the biophysical neural networks helped us to recognize the same signatures in complex interpreters of the deeper neural networks. We demonstrate how combining biophysical and deep neural networks maximizes both prediction accuracy and interpretability to yield insights into biological mechanisms.

systems biology↗

Neighborhood nonnegative matrix factorization identifies patterns and spatially-variable genes in large-scale spatial transcriptomics data

Tissues consist of multi-cellular neighborhoods in which different cell types express correlated gene programs due to shared signaling environments. Methods for identifying these spatial neighborhoods may be powerful, but currently do not scale to existing data sets of millions of cells and often artificially divide tissues into distinct neighborhoods with hard borders. To better identify multi-cellular microenvironments with shared gene programs in large-scale spatial genomics data, we developed a method that combines nonnegative matrix factorization (NMF) with Gaussian smoothing across cells in space. Our spatially-aware dimension reduction, neighborhood NMF (NNMF), identifies known and unknown interactions among diverse cell types organized into complex patterns, from localized structures to broad tissue regions. NNMF has many advantages over currently available methods, including the ability to run on modern large-scale data with thousands of features and multiple tissue samples and with millions of cells. Furthermore, our method is based on probabilistic NMF, which produces soft clusters of landscape signatures that can be understood as overlapping spatially-organized multicellular gene expression programs, allowing more biologically-complete interpretations than overly-simplistic hard clustering methods. In a benchmark dataset of a diverse set of spatial gene expression data with expert tissue labels, compared against related methods, NNMF with K-nearest neighbors clustering shows excellent performance even on hard clustering tasks. On MERFISH human colorectal cancer data, NNMF identifies several immunologically relevant multicellular interaction networks and scales to these data sets with million of cells. NNMF is implemented as an R package available at https://github.com/ragnhildlaursen/NNMF.

genomics↗