Search bioRxiv⌕ Search

Biology subjects

LeBlanc, C. J.

Publications and source records attributed to LeBlanc, C. J..

2 recordsLinked to original sources

Interpretable biophysical neural networks of transcriptional activation domains separate roles of protein abundance and coactivator binding

Deep neural networks have improved the accuracy of many difficult prediction tasks in biology, but it remains challenging to interpret these networks and learn molecular mechanisms. Here, we address the interpretability challenges associated with predicting transcriptional activation domains from protein sequence. Activation domains, regions within transcription factors that drive gene expression, were traditionally difficult to predict due to their sequence diversity and poor conservation. Multiple deep neural networks can now accurately predict activation domains, but these predictors are difficult to interpret. With the goal of interpretability, we designed simple neural networks that incorporated biophysical models of activation domains. The simplicity of these neural networks allowed us to visualize their parameters and directly interpret what the networks learned. The biophysical neural networks revealed two new ways that arrangement (i.e. the sequence grammar) of activation domain controlled function: 1) hydrophobic residues both increase activation domain strength and decrease protein abundance, and 2) acidic residues control both activation domain strength and protein abundance. Notably, the biophysical neural networks helped us to recognize the same signatures in complex interpreters of the deeper neural networks. We demonstrate how combining biophysical and deep neural networks maximizes both prediction accuracy and interpretability to yield insights into biological mechanisms.

systems biology↗

Conservation of function without conservation of amino acid sequence in intrinsically disordered transcriptional activation domains

Protein function is canonically believed to be more conserved than amino acid sequence, but this idea is only well supported in folded domains, where highly diverged sequences can fold into equivalent 3D structures with identical function. Intrinsically disordered protein regions (IDRs) often experience rapid amino acid sequence divergence, but because they do not fold into stable 3D structures, it remains unknown when and how function is conserved. As a model system for studying the evolution of IDRs, we examined transcriptional activation domains, the regions of transcription factors that bind to coactivator complexes. We systematically identified activation domains on 502 homologs of the transcriptional activator Gcn4 spanning 600 MY of fungal evolution in the Ascomycota. We find that the central activation domain shows strong conservation of function without conservation of sequence. We identify the molecular mechanism for this conservation of function without conservation of sequence: evolutionary turnover (gain and loss) of acidic and aromatic residues that are important for function. We further see turnover of complete N-terminal activation domains. This turnover at two length scales confounds multiple sequence alignments, explaining why traditional comparative genomics cannot detect functional conservation of activation domains. Evolutionary turnover of key residues is likely a general mechanism for conservation of function without conservation of sequence in IDRs.

systems biology↗